Por favor, use este identificador para citar o enlazar este ítem: http://www.alice.cnptia.embrapa.br/alice/handle/doc/1188833
Registro completo de metadatos
Campo DCValorLengua/Idioma
dc.contributor.authorOMAGE, F. B.
dc.contributor.authorNESHICH, G.
dc.date.accessioned2026-08-04T10:53:38Z-
dc.date.available2026-08-04T10:53:38Z-
dc.date.created2026-08-03
dc.date.issued2026
dc.identifier.citationGigaScience, 2026.
dc.identifier.issn2047-217X
dc.identifier.urihttp://www.alice.cnptia.embrapa.br/alice/handle/doc/1188833-
dc.descriptionBackground: Membrane proteins constitute approximately 20–30% of all proteomes and represent over 60% of current drug targets. Although protein–lipid interactions play important structural and regulatory roles in membrane-associated proteins, most existing structural resources focus on identifying whether a residue lies within a membrane region, typically inferred from computational hydrophobicity-based positioning algorithms. This approach does not directly address a distinct biological question: which residues at the protein surface make direct physical contact with lipid molecules? Answering this question from experimental data is critical for understanding lipid-mediated allostery, designing lipid-mimetic therapeutics, and training accurate machine learning models for lipid binding site prediction. Findings: We present MPLID (Membrane Protein-Lipid Interaction Database), a curated residue-level dataset comprising 4,704 membrane proteins representing 813 sequence clusters at 30% identity, 8,055,325 residues, and 80,439 experimentally validated lipid contact annotations (1.00% observed positive rate). Labels are derived exclusively from crystallized lipid molecules resolved in Protein Data Bank structures using a 4.0 ˚ A all-atom heavy-atom distance cutoff. Because most native lipid interactions are lost during purification and crystallization, this observed rate represents a lower bound, and the non-contact class inevitably contains false negatives. The dataset uses a curated list of 117 candidate lipid identifiers across ten functional categories, including 90 PDB-derived ligand codes audited against the RCSB Chemical Component Dictionary and 27 CHARMM-style lipid identifiers encountered in cryo-EM depositions. These identifiers span phospholipids, cardiolipin, sphingolipids, sterols, fatty acids, glycerolipids, detergent mimetics (explicitly flagged), lipid A components, and CHARMM simulation nomenclature. To prevent data leakage, proteins are clustered at 30% sequence identity using MMseqs2, yielding 813 clusters partitioned into training (2,578), validation (1,051), and test (1,075) splits. Amino acid composition analysis reveals biologically consistent enrichment at lipid contact sites: tryptophan (1.88×), arginine (1.44×), glycine (1.36×), lysine (1.33×), and phenylalanine (1.23×) are enriched, while proline (0.51×), isoleucine (0.57×), and aspartate (0.59×) are depleted. Conclusions: MPLID addresses a distinct biological question compared to existing resources (OPM, MemBlob, BioDolphin/PLIP): identifying residues that directly contact experimentally resolved lipid molecules rather than those positioned within computationally defined membrane boundaries. With 4,704 proteins and over 8 million annotated residues, MPLID provides the scale needed for training deep learning models for lipid contact prediction, with direct applications in structure-guided drug design and membrane protein engineering. The dataset adheres to FAIR principles and is freely available under a CC0 public domain dedication. Structurally resolved contacts represent only a subset of biological protein-lipid interactions, and MPLID is intended as an experimentally grounded resource rather than a complete catalog of lipid binding sites.
dc.language.isoeng
dc.rightsopenAccess
dc.subjectInterações lipídio-proteína
dc.subjectEstrutura de proteína
dc.subjectDados para aprendizado de máquina
dc.subjectBiologia estrutural
dc.subjectValidação experimental
dc.subjectDados FAIR
dc.subjectLipid-protein interactions
dc.subjectMachine learning dataset
dc.subjectStructural biology
dc.subjectExperimental validation
dc.subjectResidue-level classification
dc.subjectFAIR data
dc.titleMPLID (Membrane Protein–Lipid Interaction Database): a large-scale experimental resource of residue-level protein–lipid contacts.
dc.typeArtigo de periódico
dc.subject.nalthesaurusMembrane proteins
dc.subject.nalthesaurusProtein structure
dc.description.notesOn-line first.
riaa.ainfo.id1188833
riaa.ainfo.lastupdate2026-08-03
dc.identifier.doi10.1093/gigascience/giag080
dc.contributor.institutionFOLORUNSHO BRIGHT OMAGE, UNIVERSITY OF OXFORD; GORAN NESIC, CNPTIA.
Aparece en las colecciones:Artigo em periódico indexado (CNPTIA)

Ficheros en este ítem:
Fichero TamañoFormato 
AP-MPLID-2026.pdf4,27 MBAdobe PDFVisualizar/Abrir

FacebookTwitterDeliciousLinkedInGoogle BookmarksMySpace