Oracle’s AI team has joined forces with Vector Space Biosciences to advance the capabilities of large language models (LLMs) and visualizations. Additional scientific collaborators in space biosciences include
NVIDIA,
Lawrence Berkeley National Laboratory (LBNL/DOE),
University of California Berkeley (UCB),
University College London (UCL),
Imperial College of London (ICL),
IFO Rome,
City of Hope, and
McGill University.
What is Language Modeling?
Language models represent the tip of the spear in AI (Artificial Intelligence) and ML (Machine Learning) today.
Languages can exist in many different forms from a sequence of DNA/RNA, a string of amino acids, a series of nucleotides or molecular sequences to musical, mathematical, software or chemical notation - any sequence of symbols is a language including human language.
Language models and their vector representations (embeddings) are at the core of recent breakthroughs in AI including AlphaFold2's ability to predict the way a protein folds based on a sequence of amino acids, ProteinMPNN's ability to design entirely new proteins from scratch, Chroma's "DALL-E 2 for Biology" and BioNeMo, NVIDIA's large language model's (LLM) ability to generate, predict and understand biological data in new ways along with the recent release of OpenAI's ChatGPT, a breakthrough language model for human dialogue and with a few others being jointly developed by Oracle and Vector Space Biosciences. Dataset provenance and security is managed via the VXV wallet-enabled API architecture supporting utility token transactions signed on-chain.
What Can be accomplished With a correlation matrix dataset based on Language Modeling?
The CMDB API enables new ways of clustering proteins, pathways, drug compounds, molecular sequences and diseases based on context-dependent known and hidden relationships based on the latest advancements in language modeling and vector representation (embeddings).
What is a context-dependent hidden relationship?
See slides 59-67 in our NASA GeneLabs AWG presentation titled "Using Scientific Data Engineering Pipelines to Accelerate Novel Discoveries in Space Biosciences"
The Correlation Matrix Dataset Builder (CMDB) REST API
A REST-based API enabling new ways of clustering context-dependent relationships, both known and hidden [ 1 2 ] between proteins, pathways, drug compounds and molecular sequences in real-time based on the latest advancements in language modeling.
Check out our demo below or get the full Jupyter Notebook from our GitHub.
Using the CMDB API to generate a correlation matrix dataset
Below is a preview of the resulting correlation matrix dataset. Rows are human proteins, columns are atrogenes and proteins related to muscle atrophy, a common stressor during human spaceflight.
Create and display a heatmap with High Charts
Below is a preview of the resulting correlation matrix dataset heatmap visualization. Rows are human proteins, columns are atrogenes and proteins related to muscle atrophy, a common stressor during human spaceflight.
Click and drag on the heatmap to zoom in.
The Protein-Protein Interaction Network (PPIN) REST API
A REST-based API which can be used to generate a multi-level graph network from a correlation matrix dataset. The graph network represents context-dependent known and hidden relationships between proteins, pathways, drug compounds and molecular sequences.
Check out our demo below or get the full Jupyter Notebook from our GitHub.
Create and Display a Graph Network with the PPIN API using High Charts
Below is a relationship network visualization based on context-dependent known (in teal) and hidden (in yellow) relationships between atrogenes and proteins related to muscle atrophy, a common stressor during human spaceflight.
Partners and collaborators









What can be done with the CMDB API
Rows contain human proteins. Columns contain scores that represent known and hidden relationships between proteins and compounds. Vector Space Biosciences develops language models enabling the construction, visualization and interpretation of real-time hidden relationship networks between proteins, pathways, drug compounds, molecular sequences and other biochemicals to advance computational biology associated to space biosciences - the science of protecting and repairing the human body during spaceflight.
This leads to the development of countermeasures against diseases associated with stressors during human spaceflight. Countermeasures double as new forms of therapeutic applications and precision medicine for everyone on Earth. There are about 1500 new peer reviewed biomedical papers published every 24 hours at the National Library of Medicine's PubMed database. This means there’s a need for real-time language modeling and correlation matrix dataset generation which provide up to date known and hidden relationships.
FAQ
See slides 59-61 in our NASA GeneLabs AWG presentation titled "Using Scientific Data Engineering Pipelines to Accelerate Novel Discoveries in Space Biosciences"
Scores range from 0 to 1 and represent strength of known and hidden relationships between proteins, compounds, biochemicals and other biological concepts. The score is calculated based on distance calculations between vector representations (embeddings) from an ensemble of large and small language models which operate on a variety of data including molecular sequences and biomedical ontology.
- Customer proprietary databases
- Protein databases
- Molecular sequence databases
- Annotations
- Drugbank.ca
- ClinicalTrials.gov
- Ontologies
- Encyclopedias
- Dictionaries
- PubMed
- New data sources are added on a regular basis
Yes, contact us via support@vectorspacebio.science for custom configurations.
- NVIDIA BioNeMo
- BioClinicalBERT
- RoBERTa-MIMIC
- MedBERT
- AlphaBERT
- BioBERT
- BioNeMo
- Experimental (Lawrence Berkeley National Laboratory)
- New language models are added on a regular basis
- Euclidean
- Pearson
- Uncentered Pearson
- Manhattan
- EMBL
- US National Library of Medicine (NLM)
- OMIM
- UniProt
- Reactome
- Gene Ontology
- PharmaGKB
- DrugBank
- Drugs Product Database (DPD)
- PubChem Substance
- KEGG Drug
- GenBank
- Therapeutic Targets Database
- ChEMBL
- FDA label
- Product name
- Description
- Pharmacodynamics
- Molecular Weight
- Molecular Formula
- Mechanism of Action (MOA)
- Genes/Proteins
- Genomic Pathways
- Associated FASTA sequence
- Melting Point
- Hydrophobicity
- Isoelectric Point
- Metabolism
- Dosage form
- Dosage strength
- Absorption
- Drug-to-drug interactions
- Indications
- Adverse effects
- Toxicity descriptions
- Start/End marketing dates
- Half-life
- Route of elimination
- Synonyms
- Pubmed IDs/abstracts/papers
- Book reference IDs
- Generic: true/false
- Approved: yes/no
- Country
- Manufacturer
- MESH ID
- Ontology
- Encyclopedia of Biological Chemistry - 1st & 2nd Edition
- Encyclopedia of Genetics, Genomics, Proteomics and Bioinformatics
- Encyclopedia of Molecular Cell Biology and Molecular Medicine Vol 1-16
- The Encyclopedia of Molecular Biology - Creighton
- The Gale Encyclopedia of Medicine Vol 1-6 4th Edition
- Van Nostrand's Encyclopedia of Chemistry - 5th Edition
- Encyclopedia of Earth And Space Science
- Encyclopedia of Plant and Crop Science - 1st Edition
- Encyclopedia of Earth Science
- Encyclopedia of Solid Earth Geophysics
- Encyclopedia of Marine Science
- Encyclopedia of Physical Science Volume 1 & 2
- Encyclopedia of Business and Finance Vol 1 & 2
- SAGE Publications Encyclopedia of Business in Today's World
- The Encyclopedia of Political Science Set - 1st Edition
- Encyclopedia of World Geography
- Encyclopedia of Geology - 5 Volume Set
- Encyclopedia of World History 7 Volumes Set Facts on File
- Encyclopedia of Mathematical Physics Vol 1-5
- Encyclopedia of Mathematics Science
- Encyclopedia of Condensed Matter Physics
- Encyclopedia of Physics Research
- Encyclopedia of Nonlinear Science
- Rourke's World of Science Encyclopedia - Vol 1-10
- McGraw-Hill Encyclopedia of Science & Technology, 10th Edition, Vol 1-19
- Wiley Encyclopedia of Computer Science and Engineering - Vol 1-5
- Wiley Encyclopedia of Food Science and Technology - Vol 1-4
Additional references can be found here.
Request a dataset with custom features