TY - JOUR
T1 - A knowledge-driven protocol for prediction of proteins of interest with an emphasis on biosynthetic pathways
AU - Joshi, Adwait G.
AU - Harini, K.
AU - Meenakshi, Iyer
AU - Shafi, K. Mohamed
AU - Pasha, Shaik Naseer
AU - Mahita, Jarjapu
AU - Sajeevan, Radha Sivarajan
AU - Karpe, Snehal D.
AU - Ghosh, Pritha
AU - Nitish, Sathyanarayanan
AU - Gandhimathi, A.
AU - Mathew, Oommen K.
AU - Prasanna, Subramanian Hari
AU - Malini, Manoharan
AU - Mutt, Eshita
AU - Naika, Mahantesha
AU - Ravooru, Nithin
AU - Rao, Rajas M.
AU - Shingate, Prashant N.
AU - Sukhwal, Anshul
AU - Sunitha, Margaret S.
AU - Upadhyay, Atul K.
AU - Vinekar, Rithvik S.
AU - Sowdhamini, Ramanathan
PY - 2020
Y1 - 2020
N2 - This protocol describes a stepwise process to identify proteins of interest from a query proteome derived from NGS data. We implemented this protocol on Moringa oleifera transcriptome to identify proteins involved in secondary metabolite and vitamin biosynthesis and ion transport. This knowledge-driven protocol identifies proteins using an integrated approach involving sensitive sequence search and evolutionary relationships. We make use of functionally important residues (FIR) specific for the query protein family identified through its homologous sequences and literature. We screen protein hits based on the clustering with true homologues through phylogenetic tree reconstruction complemented with the FIR mapping. The protocol was validated for the protein hits through qRT-PCR and transcriptome quantification. Our protocol demonstrated a higher specificity as compared to other methods, particularly in distinguishing cross-family hits. This protocol was effective in transcriptome data analysis of M. oleifera as described in Pasha et al.Knowledge-driven protocol to identify secondary metabolite synthesizing protein in a highly specific manner.Use of functionally important residues for screening of true hits.Beneficial for metabolite pathway reconstruction in any (species, metagenomics) NGS data. (C) 2020 The Authors. Published by Elsevier B.V.
AB - This protocol describes a stepwise process to identify proteins of interest from a query proteome derived from NGS data. We implemented this protocol on Moringa oleifera transcriptome to identify proteins involved in secondary metabolite and vitamin biosynthesis and ion transport. This knowledge-driven protocol identifies proteins using an integrated approach involving sensitive sequence search and evolutionary relationships. We make use of functionally important residues (FIR) specific for the query protein family identified through its homologous sequences and literature. We screen protein hits based on the clustering with true homologues through phylogenetic tree reconstruction complemented with the FIR mapping. The protocol was validated for the protein hits through qRT-PCR and transcriptome quantification. Our protocol demonstrated a higher specificity as compared to other methods, particularly in distinguishing cross-family hits. This protocol was effective in transcriptome data analysis of M. oleifera as described in Pasha et al.Knowledge-driven protocol to identify secondary metabolite synthesizing protein in a highly specific manner.Use of functionally important residues for screening of true hits.Beneficial for metabolite pathway reconstruction in any (species, metagenomics) NGS data. (C) 2020 The Authors. Published by Elsevier B.V.
KW - Pathway
KW - Homology
KW - Multiple sequence alignment
KW - Functionally important residue
KW - Phylogenetic analysis
KW - Pathway
KW - Homology
KW - Multiple sequence alignment
KW - Functionally important residue
KW - Phylogenetic analysis
UR - https://res.slu.se/id/publ/120205
U2 - 10.1016/j.mex.2020.101053
DO - 10.1016/j.mex.2020.101053
M3 - Journal article
C2 - 33024710
SN - 2215-0161
VL - 7
JO - MethodsX
JF - MethodsX
M1 - 101053
ER -