Descifrando el Lenguaje de las Proteínas: Modelos de Lenguaje a Gran Escala para el Diseño y Optimización de Nuevas Macromoléculas
DOI:
https://doi.org/10.22201/dgtic.26832968e.2026.17.160Palabras clave:
modelos de lenguaje, generación de secuencias proteicas, diseño de proteínas, inteligencia artificial generativa, cómputo de alto desempeño, biología computacionalResumen
Los modelos de lenguaje de gran escala (LLMs, por sus siglas en inglés) han transformado la inteligencia artificial al demostrar que pueden aprender patrones estadísticos complejos a partir de datos secuenciales masivos. De manera análoga al lenguaje natural, las proteínas pueden entenderse como secuencias de aminoácidos cuya organización jerárquica determina su estructura y función. Este paralelismo ha impulsado el uso de LLMs para la generación de secuencias proteicas de novo, permitiendo explorar regiones del espacio de secuencias más allá de los métodos tradicionales. En esta revisión, presentamos los fundamentos de los LLMs aplicados a ciencia de proteínas, en particular, la arquitectura Transformer y el modelado autorregresivo, describiendo cómo los mecanismos de atención, las representaciones vectoriales y las estrategias de entrenamiento posibilitan la generación de proteínas residuo por residuo, de manera condicional o incondicional. Asimismo, analizamos los enfoques generativos y los criterios de validación computacional y experimental de las secuencias obtenidas. Finalmente, destacamos la importancia del supercómputo y de infraestructuras con aceleradores especializados para el entrenamiento y despliegue de estos modelos, así como otras aplicaciones relevantes en ciencia de proteínas, como la anotación funcional y la predicción estructural.
Descargas
Citas
[1] D. Listov, C. A. Goverde, B. E. Correia, and S. J. Fleishman, “Opportunities and challenges in design and optimization of protein function,” Nat. Rev. Mol. Cell Biol., vol. 25, no. 8, pp. 639–653, Aug. 2024, doi: 10.1038/s41580-024-00718-y.
[2] H. Y. Koh et al., “AI-driven protein design,” Nat. Rev. Bioeng., vol. 3, no. 12, pp. 1034–1056, Sep. 2025, doi: 10.1038/s44222-025-00349-8.
[3] K. I. Albanese, S. Barbe, S. Tagami, D. N. Woolfson, and T. Schiex, “Computational protein design,” Nat. Rev. Methods Primer, vol. 5, no. 1, p. 13, Feb. 2025, doi: 10.1038/s43586-025-00383-1.
[4] H. Khakzad, I. Igashov, A. Schneuing, C. Goverde, M. Bronstein, and B. Correia, “A new age in protein design empowered by deep learning,” Cell Syst., vol. 14, no. 11, pp. 925–939, Nov. 2023, doi: 10.1016/j.cels.2023.10.006.
[5] T. Kortemme, “De novo protein design—From new structures to programmable functions,” Cell, vol. 187, no. 3, pp. 526–544, Feb. 2024, doi: 10.1016/j.cell.2023.12.028.
[6] J. Beck and S. Romero-Romero, “Designing de novo TIM barrels: insights into stabilization, diversification, and functionalization strategies,” Biochem. Soc. Trans., vol. 54, no. 2, p. BST20253060, Feb. 2026, doi: 10.1042/BST20253060.
[7] A. Zubiaga, “Natural language processing in the era of large language models,” Front. Artif. Intell., vol. 6, p. 1350306, Jan. 2024, doi: 10.3389/frai.2023.1350306.
[8] D. Jurafsky and J. H. Martin, Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models, 3rd ed. 2026. [Online]. Available: https://web.stanford.edu/~jurafsky/slp3.
[9] M. Raza, “LLMs vs. SLMs: The Differences in Large & Small Language Models,” Splunk-Blogs. [Online]. Available: https://www.splunk.com/en_us/blog/learn/language-models-slm-vs-llm.html
[10] J. S. Lee, O. Abdin, and P. M. Kim, “Language models for protein design,” Curr. Opin. Struct. Biol., vol. 92, p. 103027, Jun. 2025, doi: 10.1016/j.sbi.2025.103027.
[11] E. Simon, K. Swanson, and J. Zou, “Language models for biological research: a primer,” Nat. Methods, vol. 21, no. 8, pp. 1422–1429, Aug. 2024, doi: 10.1038/s41592-024-02354-y.
[12] Standford University, “AI Demystified: Introduction to large language models,” Techology training. [Online]. Available: https://uit.stanford.edu/service/techtraining/ai-demystified/llm
[13] I. A. Blank, “What are large language models supposed to model?,” Trends Cogn. Sci., vol. 27, no. 11, pp. 987–989, Nov. 2023, doi: 10.1016/j.tics.2023.08.006.
[14] T. B. Brown et al., “Language Models are Few-Shot Learners,” Jul. 22, 2020, arXiv:2005.14165. doi: 10.48550.
[15] Microsoft, “How Embeddings Extend Your AI Model’s Reach,” Learn Microsoft. [Online]. Available: https://learn.microsoft.com/en-us/dotnet/ai/conceptual/embeddings
[16] Microsoft, “Understanding tokens,” Learn Microsoft. [Online]. Available: https://learn.microsoft.com/en-us/dotnet/ai/conceptual/understanding-tokens
[17] G. Zhang, C. Liu, J. Lu, S. Zhang, and L. Zhu, “The Role of AI-Driven De Novo Protein Design in the Exploration of the Protein Functional Universe,” Biology, vol. 14, no. 9, p. 1268, Sep. 2025, doi: 10.3390/biology14091268.
[18] Toloka Team, “History of LLMs: Complete Timeline & Evolution (1950-2026),” Toloka AI Blog. [Online]. Available: https://toloka.ai/blog/history-of-llms/
[19] K. D. Foote, “A Brief History of Natural Language Processing,” DATAVERSITY. [Online]. Available: https://www.dataversity.net/articles/a-brief-history-of-natural-language-processing-nlp/
[20] F. Tahir, “Deep Learning Models: CNN, RNN and Transformers,” Medium. [Online]. Available: https://medium.com/@fatima.tahir511/deep-learning-models-e07492b02bb0
[21] D. Mashru, “Comparative Analysis of CNN, RNN, LSTM, and Transformer Architectures in Deep Learning,” Educ. Adm. Theory Pract., pp. 5439–5443, Apr. 2023, doi: 10.53555/kuey.v29i4.10364.
[22] A. Vaswani et al., “Attention Is All You Need,” 2017, arXiv:1706.03762. doi: 10.48550
[23] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” 2018, arXiv:1810.04805. doi: 10.48550.
[24] A. R. K. Narasimhan, “Improving Language Understanding by Generative Pre-Training,” 2018. [Online]. Available: https://api.semanticscholar.org/CorpusID:49313245
[25] The UniProt Consortium et al., “UniProt: the Universal Protein Knowledgebase in 2025,” Nucleic Acids Res., vol. 53, no. D1, pp. D609–D617, Jan. 2025, doi: 10.1093/nar/gkae1010.
[26] OpenAI, “What Are Tokens and How to Count them?,” Help-OpenAI. [Online]. Available: https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
[27] T. B. Brown et al., “Language Models are Few-Shot Learners,” 2020, arXiv:2005.14165. doi: 10.48550.
[28] Simplifying Future Tech, “LLM Parameters and Hyperparameters Explained,” Medium. [Online]. Available: https://medium.com/@SimplifyingFutureTech/llm-parameters-and-hyperparameters-explained-2ba876dbae29
[29] IBM, “What is backpropagation?,” IBM-Topics. [Online]. Available: https://www.ibm.com/think/topics/backpropagation
[30] Z. ul Abideen, “Autoregressive Models for Natural Language Processing,” Medium. [Online]. Available: https://medium.com/@zaiinn440/autoregressive-models-for-natural-language-processing-b95e5f933e1f
[31] V. B. Parthasarathy, A. Zafar, A. Khan, and A. Shahid, “The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities,” 2024, arXiv: 2408.13296. doi: 10.48550.
[32] A. Takyar, “Optimize to actualize: The impact of hyperparameter tuning on AI,” LeewayHertz-AI Development Company. [Online]. Available: https://www.leewayhertz.com/hyperparameter-tuning/
[33] A. Babu, “A Comprehensive Guide to Hyperparameter Tuning in Machine Learning,” Medium. [Online]. Available: https://medium.com/@aditib259/a-comprehensive-guide-to-hyperparameter-tuning-in-machine-learning-dd9bb8072d02
[34] S. Romero-Romero, S. Lindner, and N. Ferruz, “Exploring the Protein Sequence Space with Global Generative Models,” Cold Spring Harb. Perspect. Biol., vol. 15, no. 11, p. a041471, Nov. 2023, doi: 10.1101/cshperspect.a041471.
[35] N. Ferruz, S. Schmidt, and B. Höcker, “ProtGPT2 is a deep unsupervised language model for protein design,” Nat. Commun., vol. 13, no. 1, p. 4348, Jul. 2022, doi: 10.1038/s41467-022-32007-7.
[36] E. Nijkamp, J. Ruffolo, E. N. Weinstein, N. Naik, and A. Madani, “ProGen2: Exploring the Boundaries of Protein Language Models,” 2022, arXiv:2206.13517. doi: 10.48550.
[37] A. Bhatnagar et al., “Scaling Unlocks Broader Generation and Deeper Functional Understanding of Proteins,” Apr. 16, 2025, Synthetic Biology. doi: 10.1101/2025.04.15.649055.
[38] A. Madani et al., “Large language models generate functional protein sequences across diverse families,” Nat. Biotechnol., vol. 41, no. 8, pp. 1099–1106, Aug. 2023, doi: 10.1038/s41587-022-01618-2.
[39] G. Munsamy et al., “Conditional language models enable the efficient design of proficient enzymes,” Bioinformatics, May 05, 2024, doi: 10.1101/2024.05.03.592223.
[40] T. F. Truong and T. Bepler, “PoET: A generative model of protein families as sequences-of-sequences,” 2023, arXiv: 2306.06156. doi: 10.48550.
[41] L. Lv et al., “ProLLaMA: A Protein Large Language Model for Multi-Task Protein Language Processing,” 2024, arXiv:2402.16445. doi: 10.48550/ARXIV.2402.16445.
[42] R. W. Shuai, J. A. Ruffolo, and J. J. Gray, “IgLM: Infilling language modeling for antibody sequence design,” Cell Syst., vol. 14, no. 11, pp. 979-989.e4, Nov. 2023, doi: 10.1016/j.cels.2023.10.001.
[43] E. Nguyen et al., “Sequence modeling and design from molecular to genome scale with Evo,” Science, vol. 386, no. 6723, p. eado9336, Nov. 2024, doi: 10.1126/science.ado9336.
[44] G. Brixi et al., “Genome modeling and design across all domains of life with Evo 2,” Genomics, Feb. 21, 2025, doi: 10.1101/2025.02.18.638918.
[45] T. F. Truong and T. Bepler, “Understanding protein function with a multimodal retrieval-augmented foundation model,” 2025, arXiv:2508.04724. doi: 10.48550.
[46] M. Heinzinger et al., “Bilingual language model for protein sequence and structure,” NAR Genomics and Bioinformatics, vol. 6, no. 4, p. lqae150, Nov. 2024, doi: 10.1093/nargab/lqae150.
[47] H. He et al., “De novo generation of SARS-CoV-2 antibody CDRH3 with a pre-trained generative large language model,” Nat. Commun., vol. 15, no. 1, p. 6867, Aug. 2024, doi: 10.1038/s41467-024-50903-y.
[48] X. Wang, Z. Zheng, F. Ye, D. Xue, S. Huang, and Q. Gu, “Diffusion Language Models Are Versatile Protein Learners,” 2024, arXiv:.2402.18567. doi: 10.48550.
[49] S. Romero-Romero, A. E. Braun, T. Kossendey, N. Ferruz, S. Schmidt, and B. Höcker, “De novo design of triosephosphate isomerases using generative language models,” bioRxiv Nov. 10, 2024, doi: 10.1101/2024.11.10.622869.
[50] Z. Lin et al., “Evolutionary-scale prediction of atomic-level protein structure with a language model,” Science, vol. 379, no. 6637, pp. 1123–1130, Mar. 2023, doi: 10.1126/science.ade2574.
[51] P. Notin et al., “ProteinGym: Large-Scale Benchmarks for Protein Design and Fitness Prediction,” bioRxiv, Dec. 08, 2023, doi: 10.1101/2023.12.07.570727.
[52] R. Rao et al., “Evaluating Protein Transfer Learning with TAPE,” Adv. Neural Inf. Process. Syst., vol. 32, pp. 9689–9701, Dec. 2019.
[53] F. L. Bronnec, A. Verine, B. Negrevergne, Y. Chevaleyre, and A. Allauzen, “Exploring Precision and Recall to assess the quality and diversity of LLMs,” Jun. 04, 2024, arXiv:2402.10693. doi: 10.48550/arXiv.2402.10693.
[54] J. Hon et al., “SoluProt: prediction of soluble protein expression in Escherichia coli,” Bioinformatics, vol. 37, no. 1, pp. 23–28, Apr. 2021, doi: 10.1093/bioinformatics/btaa1102.
[55] O. Conchillo-Solé, N. S. de Groot, F. X. Avilés, J. Vendrell, X. Daura, and S. Ventura, “AGGRESCAN: a server for the prediction and evaluation of ‘hot spots’ of aggregation in polypeptides,” BMC Bioinformatics, vol. 8, p. 65, Feb. 2007, doi: 10.1186/1471-2105-8-65.
[56] J. Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature, vol. 596, no. 7873, pp. 583–589, Aug. 2021, doi: 10.1038/s41586-021-03819-2.
[57] M. Leclercq and A. Droit, “Protein Language Models: Applications and Perspectives,” J. Proteome Res., Dec. 2025, doi: 10.1021/acs.jproteome.5c00506.
[58] N. Ghosh, D. Santoni, D. Nawn, E. Ottaviani, and G. Felici, “A Comprehensive Review of Transformer-based language models for Protein Sequence Analysis and Design,” 2025, arXiv:2507.13646. doi: 10.48550.
[59] J.-Y. Chen, J.-F. Wang, Y. Hu, X.-H. Li, Y.-R. Qian, and C.-L. Song, “Evaluating the advancements in protein language models for encoding strategies in protein function prediction: a comprehensive review,” Front. Bioeng. Biotechnol., vol. 13, p. 1506508, Jan. 2025, doi: 10.3389/fbioe.2025.1506508.
[60] Y. Liu and B. Tian, “Protein–DNA binding sites prediction based on pre-trained protein language model and contrastive learning,” Brief. Bioinform., vol. 25, no. 1, p. bbad488, Nov. 2023, doi: 10.1093/bib/bbad488.
[61] N. Brandes, D. Ofer, Y. Peleg, N. Rappoport, and M. Linial, “ProteinBERT: a universal deep-learning model of protein sequence and function,” Bioinformatics, vol. 38, no. 8, pp. 2102–2110, Apr. 2022, doi: 10.1093/bioinformatics/btac020.
[62] V. Thumuluri, H.-M. Martiny, J. J. Almagro Armenteros, J. Salomon, H. Nielsen, and A. R. Johansen, “NetSolP: predicting protein solubility in Escherichia coli using language models,” Bioinformatics, vol. 38, no. 4, pp. 941–946, Jan. 2022, doi: 10.1093/bioinformatics/btab801.
[63] P. Shrestha, J. Kandel, H. Tayara, and K. T. Chong, “Post-translational modification prediction via prompt-based fine-tuning of a GPT-2 model,” Nat. Commun., vol. 15, no. 1, p. 6699, Aug. 2024, doi: 10.1038/s41467-024-51071-9.
[64] K. M. Tolle, D. S. W. Tansley, and A. J. G. Hey, “The Fourth Paradigm: Data-Intensive Scientific Discovery [Point of View],” Proc. IEEE, vol. 99, no. 8, pp. 1334–1337, Aug. 2011, doi: 10.1109/JPROC.2011.2155130.
[65] J. Dongarra et al., “The International Exascale Software Project roadmap,” Int. J. High Perform. Comput. Appl., vol. 25, no. 1, pp. 3–60, Feb. 2011, doi: 10.1177/1094342010391989.
[66] C. T. Lee and R. E. Amaro, “Exascale Computing: A New Dawn for Computational Biology,” Comput. Sci. Eng., vol. 20, no. 5, pp. 18–25, Sep. 2018, doi: 10.1109/MCSE.2018.05329812.
[67] Big Data Interagency Working Group, High End Computing Interagency Working Group, Networking & Information Technology Research & Development Subcommittee, and Committee On Science & Technology Enterprise Of The National Science & Technology Council, The Convergence of High Performance Computing, Big Data, and Machine Learning. U.S Government, Sep. 2019. [Online]. Available: https://www.nitrd.gov/pubs/Convergence-HPC-BD-ML-JointWSreport-2019.pdf
[68] R. Bommasani et al., “On the Opportunities and Risks of Foundation Models,” Jul. 12, 2022, arXiv:2108.07258. doi: 10.48550.
[69] UNESCO, “Implementation of the UNESCO Recommendation on Open Science,” UNESCO - Open Science. [Online]. Available: https://www.unesco.org/en/open-science/implementation
[70] A. Santillán González and L. Hernández-Cervantes, “Hitos del supercómputo: del Giga al Exascale,” EPISTEMUS, vol. 18, no. 35, Dec. 2023, doi: 10.36790/epistemus.v18i35.300.
[71] S. Ren, B. Tomlinson, R. W. Black, and A. W. Torrance, “Reconciling the contrasting narratives on the environmental impact of large language models,” Sci. Rep., vol. 14, no. 1, p. 26310, Nov. 2024, doi: 10.1038/s41598-024-76682-6.
[72] Y. Shen, G. Kudla, and D. A. Oyarzún, “Improving the generalization of protein expression models with mechanistic sequence information,” Nucleic Acids Res., vol. 53, no. 3, p. gkaf020, Jan. 2025, doi: 10.1093/nar/gkaf020.
[73] P. Hunter, “Security challenges by AI-assisted protein design: The ability to design proteins in silico could pose a new threat for biosecurity and biosafety,” EMBO Rep., vol. 25, no. 5, pp. 2168–2171, Mar. 2024, doi: 10.1038/s44319-024-00124-7.
[74] X. Fang et al., “A method for multiple-sequence-alignment-free protein structure prediction using a protein language model,” Nat. Mach. Intell., vol. 5, no. 10, pp. 1087–1096, Oct. 2023, doi: 10.1038/s42256-023-00721-6.
[75] H. Ghazikhani and G. Butler, “TooT-BERT-M: Discriminating Membrane Proteins from Non-Membrane Proteins using a BERT Representation of Protein Primary Sequences,” in 2022 IEEE Conference on Computational Intelligence in Bioinformatics and Computational Biology (CIBCB), Ottawa, ON, Canada: IEEE, Aug. 2022, pp. 1–8. doi: 10.1109/CIBCB55180.2022.9863026.
[76] N. Buton, F. Coste, and Y. Le Cunff, “Predicting enzymatic function of protein sequences with attention,” Bioinformatics, vol. 39, no. 10, p. btad620, Oct. 2023, doi: 10.1093/bioinformatics/btad620.
[77] Q. Liu, C. Zhang, and L. Freddolino, “InterLabelGO+: unraveling label correlations in protein function prediction,” Bioinformatics, vol. 40, no. 11, p. btae655, Nov. 2024, doi: 10.1093/bioinformatics/btae655.
[78] G. Li, S. Yao, and L. Fan, “ProSTAGE: Predicting Effects of Mutations on Protein Stability by Using Protein Embeddings and Graph Convolutional Networks,” J. Chem. Inf. Model., vol. 64, no. 2, pp. 340–347, Jan. 2024, doi: 10.1021/acs.jcim.3c01697.
[79] Y. Luo, Z. Nie, M. Hong, S. Zhao, H. Zhou, and Z. Nie, “MutaPLM: Protein Language Modeling for Mutation Explanation and Engineering,” 2024, arXiv:2410.22949. doi: 10.48550.
[80] F. Teufel et al., “SignalP 6.0 predicts all five types of signal peptides using protein language models,” Nat. Biotechnol., vol. 40, no. 7, pp. 1023–1025, Jul. 2022, doi: 10.1038/s41587-021-01156-3.
[81] H. K. Wayment-Steele et al., “Learning millisecond protein dynamics from what is missing in NMR spectra,” bioRxiv, Mar. 19, 2025, doi: 10.1101/2025.03.19.642801.
[82] W. Liu et al., “PLMSearch and PLMAlign: Protein Language Model (PLM)-Based Homologous Protein Sequence Search and Alignment,” in Large Language Models (LLMs) in Protein Bioinformatics, vol. 2941, D. B. Kc, Ed., in Methods in Molecular Biology, vol. 2941. , New York, NY: Springer US, 2025, pp. 227–241. doi: 10.1007/978-1-0716-4623-6_14.
[83] U. Lupo, D. Sgarbossa, and A.-F. Bitbol, “Pairing interacting protein sequences using masked language modeling,” Proc. Natl. Acad. Sci., vol. 121, no. 27, p. e2311887121, Jul. 2024, doi: 10.1073/pnas.2311887121.
[84] S. J. Giri, N. Ibtehaz, and D. Kihara, “GO2Sum: generating human-readable functional summary of proteins from GO terms,” Npj Syst. Biol. Appl., vol. 10, no. 1, p. 29, Mar. 2024, doi: 10.1038/s41540-024-00358-0.
Descargas
Publicado
Cómo citar
Número
Sección
Licencia
Derechos de autor 2026 Sergio Romero-Romero, Diego Pérez-Villanueva, Luz Mariana Gonzalez-Vega, Diego Muñoz Beltran

Esta obra está bajo una licencia internacional Creative Commons Atribución-NoComercial 4.0.
TIES, Revista de Tecnología e Innovación en Educación Superior, es una publicación semestral de acceso abierto bajo la licencia Creative Commons Atribución-No Comercial 4.0 Internacional (CC BY-NC 4.0).
ISSN 22683-2968 • © 2026 Universidad Nacional Autónoma de México. TIES, Revista de Tecnología e Innovación en Educación Superior es editada por la Universidad Nacional Autónoma de México a través de la Dirección General de Cómputo y de Tecnologías de Información y Comunicación (DGTIC). Circuito exterior s/n, Ciudad Universitaria, Alcaldía Coyoacán, C.P. 04510, Ciudad de México, México • Reserva de Derechos de Autor otorgado por INDAUTOR: 04-2019-011816190900-203.
El contenido de los artículos es responsabilidad de los autores y no refleja el punto de vista del Comité editorial, del Editor o de la Universidad Nacional Autónoma de México. Hecho en México, 2026.
