String analysis and algorithms with genomic applications

dc.contributor.advisorChairperson, Graduate Committee: Binhai Zhuen
dc.contributor.authorLiyana Ralalage, Adiesha Lakshan Liyanageen
dc.date.accessioned2024-11-09T17:44:33Z
dc.date.issued2024en
dc.description.abstractIn biology, genome rearrangements are mutations that change the gene content of a genome or the arrangement of the genes on a genome. Understanding how genome rearrangements occur in a genome can help us to understand the evolutionary history of extant species, improve genetic engineering, and understand the basis of genetic diseases. In this dissertation, we explored four problems related to genome partitioning and tandem duplication and deletion rearrangement operations. Our interest was focused on determining how difficult it is to solve these problems and identifying efficient algorithms to solve them. The proposed problems were formulated as string problems and then analyzed using complexity theory. In the first chapter, we explored several variations of F -strip recovery problem called XSR-F and GSR-F and their complexity under different parameters. We proved that the XSR-F problem is hard to solve unless we restrict the allowed block sizes to one size. We provided a polynomial time algorithm for GSR-F under a fixed alphabet and fixed F . In the second and third chapters, we introduced two string problems named longest letter- duplicated subsequence (LLDS) and longest subsequence-repeated subsequence (LSRS)-- formulated as alternative problem formulations for the tandem-duplication distance problem that allow to extract information about segments of genes that may have undergone tandem duplication-- analyzed the complexity of their variations and devised efficient algorithms to solve them. We proved that constrained versions of LLDS and LSRS problems are NP- hard for parameter d > or = 4, while general versions were polynomially solvable which hints that any variations closer to the original tandem duplication distance problem are still hard to solve. In the final chapter, we delved into two heuristic algorithms designed to compute genomic distance between two mitochondrial genomes and a heuristic algorithm to predict ancestral gene order under the TDRL (tandem-duplication random loss) model. We improved the previously studied method developed for permutation strings by tweaking heuristic choices aimed at calculating the minimum distance between two genomes to apply to non-permutation strings. These heuristic algorithms were implemented and tested on a real-world mitochondrial genome data set.en
dc.identifier.urihttps://scholarworks.montana.edu/handle/1/18538
dc.language.isoenen
dc.publisherMontana State University - Bozeman, College of Engineeringen
dc.rights.holderCopyright 2024 by Adiesha Lakshan Liyanage Liyana Ralalageen
dc.subject.lcshGenomicsen
dc.subject.lcshProtein-protein interactionsen
dc.subject.lcshComputational complexityen
dc.subject.lcshAlgorithmsen
dc.titleString analysis and algorithms with genomic applicationsen
dc.typeDissertationen
mus.data.thumbpage107en
thesis.degree.committeemembersMembers, Graduate Committee: Brendan Mumey; Lucia Williams; Sean Yawen
thesis.degree.departmentComputingen
thesis.degree.genreDissertationen
thesis.degree.namePhDen
thesis.format.extentfirstpage1en
thesis.format.extentlastpage134en

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
liyana-ralalage-string-2024.pdf
Size:
992.14 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
825 B
Format:
Plain Text
Description: