Research Projects
At the Language Biomarker Lab (LaBL), we explore how machine learning and language analysis can uncover insights into human cognition and mental health.
Our research focuses on how Natural Language Processing (NLP) and Artificial Intelligence (AI) can, in essence, data mine the mind—revealing patterns in language that reflect thought, behavior, and neurological health.
We currently use advanced deep learning models—such as LLaMA 4, FLAN-UL2, and the latest versions of GPT—to investigate decision-making, the interplay between semantics and syntax, and the early markers of brain-related illnesses like Alzheimer's disease. Our work is supported by the National Institutes of Health, including the NIMH and NIA.

Linguistic Biomarkers and AI: Uncovering the Language of Mental Health
Dr. Phillip Wolff leads research at the forefront of using machine learning and natural language processing to identify linguistic markers of mental illness. His work includes developing algorithms to measure semantic density, analyze latent content, and decode time-related language. These methods have been applied to diverse contexts—from Reddit posts to transformer models like T5—to reveal how subtle patterns in language reflect cognitive and emotional states. Through this research, Dr. Wolff is helping to redefine how we detect and understand mental health conditions through the lens of language.

Causal Cognition: From Language to Perception
LaBL's research on causal cognition spans three core areas: how causal meaning is encoded in language, how causation is perceived from sensory experience, and how people reason about causal relationships. Over the past decade, Dr. Wolff has advanced the idea that causal understanding is grounded in force dynamics—reflected both in linguistic semantics and neural processes. Using methods such as computational modeling, computer visualization, haptic rendering, and corpus analysis, his work reveals the deep interconnection between language, perception, and reasoning.

Language and Thought: Tracing the Cognitive Impact of Words
LaBL's early work on linguistic biomarkers stems from a broader investigation into the relationship between language and thought. A central focus of this research has been how word meanings reveal the structure of conceptual representations. Co-editing a foundational volume with Dr. Barbara Malt on the language-thought interface, Dr. Wolff explored whether the language we speak shapes how we think—addressed through cross-linguistic studies on the encoding of causation. His research demonstrates how subtle linguistic differences across English, Korean, Chinese, and Russian can influence cognitive processes, including event perception and categorical thinking.
Featured Projects
ProNET
Psychosis Risk Outcomes Network
A global effort to better understand, treat, and prevent psychosis. ProNET is funded by the NIH and the Foundation for the NIH (FNIH) as an Accelerating Medicines Partnership (AMP), a program that brings together NIH, private foundations, and industry sponsors to accelerate the development of new medications through large-scale research initiatives. The target population includes youth and young adults at Clinical High Risk for Psychosis (CHR-P), identified through the Positive Symptoms and Diagnostic Criteria for the CAARMS Harmonized with the SIPS (PSYCHS)—a new assessment tool integrating the two most widely used psychosis risk instruments, the CAARMS and the SIPS.
Learn more →ProCan
Proof-of-Principle Clinical Trial
ProCan is the proof-of-principle clinical trial network within the Accelerating Medicines Partnership in Schizophrenia (AMP SCZ). It builds on the observational work of ProNET and PRESCIENT by testing targeted interventions in individuals at clinical high risk for psychosis, aiming to translate biomarker findings—including language-based markers—into actionable clinical tools.
Learn more →ANNA-PR
Automated Neural Network Assessment of Psychosis Risk
An AI-driven system designed to automate the assessment of psychosis risk. ANNA-PR leverages neural network architectures and natural language processing to analyze clinical interviews and language samples, enabling scalable, consistent, and objective psychosis risk evaluation.
Semantic Spectroscopy
Decomposing Meaning in Clinical Language
Semantic Spectroscopy is a computational framework for decomposing the meaning of language into its underlying semantic dimensions—much like light split through a prism. Applied to clinical speech and text data, this approach identifies fine-grained shifts in semantic content, coherence, and specificity that may serve as early indicators of psychosis, cognitive decline, and other neurological conditions.
Theme Clustering for CHR Prediction
Predicting Clinical High Risk Using AMP SCZ Data
This project applies unsupervised topic modeling and clustering methods to large-scale language data collected across the AMP SCZ network to identify recurring thematic patterns in the speech of individuals at clinical high risk for psychosis (CHR). By uncovering latent themes in clinical interviews, the work aims to build predictive models that can distinguish CHR individuals from healthy controls and track symptom progression over time.