Early detection and risk prediction system for schizophrenia

By collecting and integrating voice, behavior and clinical data in real time through a multimodal attention mechanism, visual risk prediction results are generated, which solves the lag problem in early detection of schizophrenia and realizes dynamic risk assessment and explainable early warning.

CN120766973AInactive Publication Date: 2025-10-10ZHOUSHAN SECOND PEOPLES HOSPITAL
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510958811.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2025-10-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies in the early identification of schizophrenia have problems such as screening lag, lack of dynamic risk assessment, modality splitting and model uninterpretability, resulting in high rates of missed detection or false alarms.

Method used

Multimodal data including speech, behavior and clinical assessment are collected in real time through mobile terminals. The cross-modal attention mechanism is used to dynamically weight and fuse features to generate fused feature vectors and input them into the machine learning classifier to achieve dynamic risk assessment. The classifier weights are updated online through incremental learning to generate visual risk trajectory maps and feature contribution heat maps.

Benefits of technology

It achieves real-time and explainable early detection of schizophrenia, dynamically tracks the symptom evolution of high-risk individuals, provides reliable risk predictions and assists clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766973A_ABST
    Figure CN120766973A_ABST
Patent Text Reader

Abstract

The invention relates to a system for early detection and risk prediction of schizophrenia. The system comprises a feature extraction module which performs parallel processing on a multi-mode original signal; the fusion prediction module fuses the voice feature signal, the behavior feature signal and the clinical feature signal through a cross-modal attention weighting mechanism to generate a fusion feature vector, and outputs an initial risk prediction signal through a machine learning classifier based on the vector; the dynamic updating module receives an incremental characteristic signal triggered by newly added collected data of a user, generates an updating risk prediction signal through incremental learning, and feeds back the updating risk prediction signal to the fusion prediction module to optimize the weight of a classifier; and the visual output module converts the updated risk prediction signal into a dynamic risk trajectory diagram and an interpretable feature contribution thermodynamic diagram. The system for early detection and risk prediction of schizophrenia can solve the problems that early screening of schizophrenia lags behind and dynamic risk assessment is missing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of early warning and artificial intelligence assisted diagnosis of mental illness, in particular to a system for early detection and risk prediction of schizophrenia. BACKGROUND

[0002] Current early identification of schizophrenia mainly relies on clinical interviews (such as SIPS scale), which requires professional physicians to spend time evaluating and is highly subjective. Although neuroimaging can provide evidence of brain structure / function abnormalities, it is expensive and difficult to monitor dynamically. In recent years, AI prediction based on a single modality has made progress: voice analysis detects semantic confusion through acoustic features (pitch perturbation, abnormal pause) or NLP technology, but is easily disturbed by environmental noise; wearable devices collect movement patterns that can reflect negative symptoms (such as social withdrawal), but it is difficult to distinguish schizophrenia from comorbidities such as depression; genetic and blood biomarker research has not yet been clinically applied due to heterogeneity issues.

[0003] Multimodal fusion has become a frontier direction, such as combining MRI and cognitive test prediction models (accuracy rate about 70-80%), but there are three major limitations: static modeling: based on single data collection, it cannot track the evolution of symptoms in high-risk individuals. Cohort studies have shown that 30% of patients at the prodromal stage have a significant change in risk status within six months, and static models lead to missed detection or false positives; modality fragmentation: existing systems often simply concatenate features (such as brain connectivity matrix + scale score), ignoring cross-modal interaction mechanisms. For example, language disorder and movement retardation may share neural circuit abnormalities, but traditional fusion cannot mine such correlations; clinical disconnection: the black box model outputs risk values lack interpretability, making it difficult for doctors to trust. The European PRONIA project attempts to provide brain region activation heat maps, but does not correlate specific behavior or language features. SUMMARY

[0004] In view of the shortcomings of the prior art, the purpose of the present application is to provide a system for early detection and risk prediction of schizophrenia, to solve the problems of lagging early screening of schizophrenia and lack of dynamic risk assessment. The present application collects multimodal data of voice, behavior and clinical assessment in natural situations through a mobile terminal in real time, dynamically weights and fuses voice acoustic and semantic features, behavior spatiotemporal pattern features and clinical structured features using a cross-modal attention mechanism after modality-specific feature extraction, generates a fusion feature vector input into a machine learning classifier to output risk prediction; based on the new data triggering incremental learning to update the classifier weights online to realize dynamic risk assessment, and converting the prediction results into visual dynamic risk trajectory and feature contribution heat maps to assist clinical decision-making, realizing a closed loop from continuous monitoring to interpretable early warning.

[0005] The present application provides a system for early detection and risk prediction of schizophrenia, comprising: A multi-source data acquisition module forms multi-modal raw signals including voice data, behavioral digital phenotype data, and clinical assessment data; A feature extraction module performs parallel processing on the multi-modal raw signals: acoustic and semantic features of the voice data are extracted to form voice feature signals, spatiotemporal pattern features of the behavioral digital phenotype data are extracted to form behavioral feature signals, and structured indicators of the clinical assessment data are extracted to form clinical feature signals; A fusion prediction module fuses the voice feature signals, the behavioral feature signals, and the clinical feature signals through a cross-modal attention weighting mechanism, generates a fusion feature vector, and outputs an initial risk prediction signal based on the vector through a machine learning classifier; A dynamic update module receives an incremental feature signal triggered by newly added user acquisition data, generates an updated risk prediction signal through incremental learning, and feeds it back to the fusion prediction module to optimize the classifier weights; A visual output module converts the updated risk prediction signal into a dynamic risk trajectory graph and an interpretable feature contribution heat map.

[0006] In an embodiment of the present application, the voice data in the multi-source data acquisition module is collected by a mobile terminal microphone to obtain real-time original audio streams of natural conversations or specified text reading, the behavioral digital phenotype data is collected by relying on the built-in sensors of the intelligent terminal to continuously monitor the user's motion trajectory, screen operation frequency, and environmental interaction time series, and the clinical assessment data is collected by using a structured electronic form to record the standardized scores of positive symptoms, negative symptoms, and cognitive function given by doctors, and the three types of data are encrypted and compressed and then packaged as multi-modal raw signals with a unified timestamp.

[0007] In an embodiment of the present application, when the feature extraction module extracts the voice feature signals, it first analyzes the fundamental frequency contour, the formant distribution, and the nonlinear dynamics parameters to generate an acoustic feature vector through a pre-trained acoustic model, and simultaneously uses a deep semantic encoder to analyze the text coherence and the degree of abnormality of the logical structure to generate a semantic feature vector, and the two are spliced and then processed by dimension reduction to form the voice feature signals; in the extraction of the behavioral feature signals, the time-frequency transform algorithm is used to separate the walking rhythm, gesture stability, and diurnal activity intensity features from the accelerometer and gyroscope raw data, and a Markov chain model is constructed to quantify the state transition probability; the clinical feature signals are mapped to a continuous vector space through an embedding layer.

[0008] In one embodiment of the present invention, the cross-modal attention weighting mechanism of the fusion prediction module is specifically as follows: the speech feature signals, behavioral feature signals and clinical feature signals are respectively projected into the common latent space, the cross-attention weight matrix between each pair of modalities is calculated, and the cross-modal interaction features are generated according to the dynamic weighted sum of the weights. After being spliced ​​with the independent features of each modality, the features are input into the gated recurrent unit for time series modeling, and the final output fusion feature vector is connected to a machine learning classifier composed of a support vector machine containing a radial basis function kernel.

[0009] In one embodiment of the present invention, the incremental learning process of the dynamic update module includes: after receiving the incremental feature signal triggered by the newly collected data, first using the same parameters of the feature extraction module to extract the incremental speech, behavior and clinical features, and then inputting them together with the historical fusion feature vector into the online sequence extreme learning machine, updating the classifier weights by iteratively solving the least squares solution of the hidden layer output weights, and at the same time using the sliding time window mechanism to eliminate expired features to control the complexity of the model, and the updated risk prediction signal is fed back to the fusion prediction module in real time to cover the old classifier.

[0010] In one embodiment of the present invention, when the visualization output module generates a dynamic risk trajectory diagram, the updated risk prediction signals at different time points are sorted by acquisition time, and a continuous risk probability curve is generated using cubic spline interpolation, and the clinical intervention event nodes and symptom deterioration threshold lines are marked on the curve; the feature contribution heat map quantifies the contribution of each modal feature to the current prediction result by calculating the Shapley plus explanation value, maps it on the modality-feature two-dimensional matrix with a color gradient, and displays the original data fragments in association for verification.

[0011] In one embodiment of the present invention, a deep semantic encoder is constructed using a multi-head self-attention mechanism. Its input sequence is the word segmentation embedding vector after speech-to-text conversion. Context dependencies are captured by stacking bidirectional long short-term memory network layers. The output layer uses a convolutional neural network to extract local semantic patterns. The final semantic feature vector is the maximum pooling result of the hidden states of each network layer.

[0012] In one embodiment of the present invention, a gated recurrent unit introduces residual connections and layer normalization operations in the temporal modeling process, and the weight matrices of the reset gate and update gate in its hidden state update function are dynamically initialized by cross-modal interaction features. The unit output is input into a support vector machine classifier after discarding regularization processing.

[0013] In one embodiment of the present invention, the hidden layer node activation function of the online sequential extreme learning machine adopts a rectified linear unit with a leakage parameter, its input layer is fully connected with the newly added feature signal and the historical fusion feature vector, the output layer error backpropagation only adjusts the weight of the last layer, and the weight matrix from the hidden layer to the output layer is iteratively updated through the Moore-Penrose generalized inverse matrix.

[0014] In one embodiment of the present application, the system for early detection and risk prediction of schizophrenia is deployed in a cloud edge collaborative architecture, the multi-source data acquisition module and the feature extraction module run on the edge device of the user's smart terminal, the fusion prediction module and the dynamic update module are deployed on the cloud server, the visual output module is realized through the webpage interactive interface, and the intermediate feature signal is transmitted between the edge device and the cloud through a secure channel based on quantum key distribution. The present application provides a system for early detection and risk prediction of schizophrenia. The system collects multi-modal data such as speech, behavior and clinical assessment in natural environment through mobile terminal in real time, extracts features by modality, dynamically weights and fuses acoustic semantic features, behavior spatiotemporal pattern features and clinical structured features by cross-modal attention mechanism, generates fusion feature vector, inputs machine learning classifier, and outputs risk prediction. The system realizes dynamic risk assessment by online updating classifier weight based on incremental learning triggered by new data, and converts prediction results into visual dynamic risk trajectory graph and feature contribution heat map to assist clinical decision making, realizing a closed loop from continuous monitoring to explainable early warning. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creating any creative labor.

[0016] Figure 1 It is a system architecture diagram of a system for early detection and risk prediction of schizophrenia. DETAILED DESCRIPTION

[0017] The embodiments of the present application will be described below through specific concrete examples. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in the present specification. The present application can also be implemented or applied by different specific embodiments, and the details in the present specification can be modified or changed based on different views and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.

[0018] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner, and the diagrams only show the components related to the present application, not the number, shape and size of the components when actually implemented. The actual implementation of each component may be a random change, and the component layout pattern may be more complex.

[0019] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0020] See Figure 1 , shown is a system for early detection and risk prediction of schizophrenia of the present invention, including a multi-source data acquisition module, which forms a multimodal original signal including voice data, behavioral digital phenotype data and clinical assessment data; a feature extraction module, which performs parallel processing on the multimodal original signal: extracting the acoustic and semantic features of the voice data to form a voice feature signal, extracting the spatiotemporal pattern features of the behavioral digital phenotype data to form a behavioral feature signal, and extracting the structured indicators of the clinical assessment data to form a clinical feature signal; a fusion prediction module, which fuses the voice feature signal, the behavioral feature signal and the clinical feature signal through a cross-modal attention weighting mechanism to generate a fusion feature vector, and outputs an initial risk prediction signal based on the vector through a machine learning classifier; a dynamic update module, which receives the incremental feature signal triggered by the user's newly collected data, generates an updated risk prediction signal through incremental learning and feeds it back to the fusion prediction module to optimize the classifier weight; a visualization output module, which converts the updated risk prediction signal into a dynamic risk trajectory map and an interpretable feature contribution heat map.

[0021] Figure 1 As shown, a closed-loop processing chain with multi-module collaboration is constructed. During system initialization, the multi-source data acquisition module activates user authorization through the smart terminal application, simultaneously turning on the microphone, motion sensor, and electronic clinical assessment interface. Voice data collection scenarios include natural conversation and standardized text reading. The former captures 60-second voice clips from users' daily communications in real time (with an ambient signal-to-noise ratio threshold set to 15dB), while the latter guides users to read emotionally neutral short passages and record the audio stream. Behavioral digital phenotyping data collection utilizes the terminal's accelerometer, gyroscope, and screen touch sensor to continuously monitor three-dimensional motion trajectories, finger sliding speed sequences, and application switching frequency at a 50Hz sampling rate. Clinical assessment data is entered via a tablet computer on the physician's end, using a structured form to record the 30-item Positive and Negative Syndrome Scale (PANSS) scores, cognitive function assessments (such as word fluency test scores), and the Checklist of Prodromal Symptoms (SOPS) results. These three types of data are encapsulated into time-aligned data packets, forming a multimodal raw signal that is transmitted to the feature extraction module.

[0022] Further, the feature extraction module is deployed on the edge computing node, adopting a split-modal parallel processing architecture: the speech data processing thread calls a pre-trained convolutional neural network to extract short-time spectrogram features, and synchronously runs a speech recognition engine to convert audio to text and then input a bidirectional long short-term memory network to extract semantic coherence indicators; the behavior data processing thread applies wavelet transform to decompose the energy distribution characteristics of the accelerometer signal, and combines a hidden Markov model to identify motion mode mutation points; the clinical data processing thread maps discrete scale scores to 128-dimensional vectors through an embedding layer. The generated speech feature signals (acoustic features 32-dimensional + semantic features 64-dimensional), behavior feature signals (time domain features 24-dimensional + frequency domain features 16-dimensional) and clinical feature signals (128-dimensional) are standardized and transmitted to the fusion prediction module. The fusion prediction module runs on a cloud server, and its cross-modal attention weighting mechanism is specifically implemented as follows: first, project the three types of feature signals into a 256-dimensional common hidden space, calculate three groups of cross-attention weight matrices for speech-behavior, speech-clinical and behavior-clinical, dynamically weight the feature vectors according to the weights, and generate 512-dimensional cross-modal interaction features. The feature is concatenated with the original modal feature and input into the gated recurrent unit network, and the time-dependent fusion feature vector is output. Finally, the initial risk prediction signal (0-1 continuous probability value) is generated through the radial basis function support vector machine classifier. This signal triggers the listening service of the dynamic update module, and when a user new data acquisition event occurs, the incremental feature signal is input into the online sequence extreme learning machine through the same feature extraction process. This algorithm takes the historical fusion feature vector as the initial state, and realizes the weight update of the classifier by iteratively solving the generalized inverse of the hidden layer weight matrix, and the updated risk prediction signal real-time covers the old model. After receiving the updated signal, the visual output module calls the Sharpe value interpretation algorithm to calculate the contribution of each modal feature, generates a dynamic risk trajectory graph (horizontal axis is time, vertical axis is risk probability) and a modal-feature two-dimensional heat map (color mapping contribution intensity), and pushes it to the doctor workstation interface through the WebSocket protocol. During system operation, all data transmission is encrypted end-to-end using the national SM4 algorithm to ensure the security of sensitive medical data.

[0023] In one embodiment of the present invention, the voice data acquisition subsystem of the multi-source data acquisition module includes an environmental perception and quality control unit. For natural conversation acquisition, a blind source separation algorithm is first used to eliminate background noise (such as keyboard tapping and ambient music), and then an endpoint detection algorithm is used to capture valid segments of human voice. For standardized reading acquisition, the user is required to read a 500-word, emotionally neutral text (such as a news summary) in a soundproof environment. The system simultaneously monitors the volume fluctuation range (65-75dB) and speech rate uniformity (3-5 words / second). Any abnormal data automatically triggers a re-recording mechanism. The collected raw audio stream is compressed using Opus encoding and stored as a WAV file with a 16kHz sampling rate and 16-bit depth, with a timestamp and device ID metadata appended. The behavioral digital phenotyping subsystem achieves multi-sensor spatiotemporal synchronization: the accelerometer and gyroscope collect three-axis linear acceleration (range ±8g) and angular velocity (range ±2000° / s) at a 50Hz frequency, calculating the device's attitude angle using a quaternion fusion algorithm. The screen touch sensor records the touch point coordinate sequence and press duration, and combines this with application usage logs to extract metrics such as social media app opening frequency and message reply latency. A dedicated coprocessor time-aligns all sensor data (to within 10ms) and generates a JSON-structured data packet containing a timestamp, sensor type, and data vector.

[0024] like Figure 1 As shown, the clinical assessment data collection subsystem utilizes a dynamic form engine. After the physician logs in and authenticates, the system loads different assessment templates based on the user's risk level (e.g., the SOPS scale plus a word memory test for high-risk individuals, and the PANSS scale plus the Wisconsin Card Sorting Test for suspected patients). The scoring interface incorporates built-in logical validation rules (e.g., a specific behavioral description is mandatory for positive symptom scores > 6). Upon submission, the timestamps of the voice and behavioral data for the current collection period are automatically linked. Data encryption utilizes a hybrid encryption scheme based on elliptic curves: the structured score sheet is encrypted using the SM3 algorithm to generate a digital summary, which is then encrypted using the SM2 public key to form a digital envelope for transmission to the cloud. When the three types of data are encapsulated into a unified data packet at the edge device, an absolute time synchronization mechanism is employed: based on the GPS timing signal, the voice data starting point, the first behavioral sensor data point, and the clinical assessment submission time are all aligned to UTC, with a maximum time deviation of less than 50ms. The data packet structure consists of a header (protocol version, data source hash value), a payload (voice WAV binary stream, behavioral JSON text, clinical assessment XML), and a trailer (cyclic redundancy check code). The data packet is transmitted to the feature extraction module via the MQTT protocol.

[0025] like Figure 1As shown, the feature extraction module generates speech feature signals using a two-path heterogeneous processing architecture: The acoustic feature extraction path: After the raw audio input, it first passes through a pre-emphasis filter (coefficient 0.97) to compensate for high-frequency attenuation. After frame processing (frame length 25ms, frame shift 10ms), the Mel-frequency cepstral coefficients (24 dimensions) are calculated for each frame. The fundamental frequency trajectory (using the PRAAT algorithm to eliminate octave errors), formant bandwidth (F1-F3), and nonlinear features (recursive quantitative analysis entropy value and detrended fluctuation analysis scaling index) are simultaneously extracted. These features are then reduced to 32 dimensions using principal component analysis to form an acoustic feature vector. The semantic feature extraction path: The speech recognition engine uses an end-to-end streaming model (based on the Conformer architecture) to output a time-stamped text sequence in real time. After text is input into the semantic analysis unit, dependency parsing is first performed to generate a syntax tree. The subtree depth variance is calculated to measure syntactic complexity. A 300-dimensional word embedding layer is then used to construct a text matrix, which is then fed into a multi-headed self-attention encoder (6 layers, 8 heads) to extract contextual representations. Finally, a one-dimensional convolutional neural network (kernel width = [3, 5, 7]) is used to capture local semantic patterns. The outputs of each convolutional layer are then concatenated into a 64-dimensional semantic feature vector after global max pooling. The extraction of behavioral feature signals focuses on quantifying spatiotemporal patterns. The raw accelerometer signal is low-pass filtered with a Butterworth filter (cutoff frequency 20Hz) to remove high-frequency noise. It is then decomposed into five scales (corresponding to the 0.5-10Hz frequency band) using a continuous wavelet transform (Morlet wavelet basis). The energy contribution and time-frequency entropy of each scale are calculated. Motion pattern recognition utilizes a hidden Markov model, defining four states: stillness, walking, running, and gestures. The Viterbi algorithm is used to decode the most likely state sequence and calculate the state transition probability matrix (4×4) and the distribution of single event durations (Gamma distribution fitting parameters). Screen touch features are extracted, including touch point movement velocity (first-order difference), pressure variance (Android devices), and the application switching matrix (inter-application transition probability), ultimately merging these into a 40-dimensional behavioral feature signal.

[0026] Furthermore, clinical feature signal generation utilizes a medical knowledge-guided embedding approach. Discrete scale scores (e.g., items P1-P7 of the PANSS) are first converted to standard scores (T-scores) using a table lookup method. These scores are then fed into a three-layer fully connected network (hidden layer dimension 64, ReLU activation) to learn a continuous vector representation. Open-ended text descriptions (e.g., physician observation notes) are extracted using the BioClinicalBERT model to extract clinical entity embeddings. Finally, all clinical indicators are concatenated into a 128-dimensional vector, batch normalized, and output. Before transmission, each modality's feature signal undergoes a feature selection module (based on random forest importance ranking) to eliminate redundant dimensions and ensure subsequent fusion efficiency. The module hardware deployment utilizes a heterogeneous computing strategy: acoustic feature extraction runs on the device's DSP chip (optimizing FFT computation), while semantic analysis is offloaded to an edge GPU server. Behavioral feature processing utilizes a sensor coprocessor (with a built-in HMM hardware accelerator). Clinical feature generation is performed on a cloud-based CPU cluster. Each processing unit executes in parallel through a pipeline mechanism, keeping the latency of a single feature extraction pass within 800ms (meeting real-time requirements).

[0027] like Figure 1As shown in the figure, the specific process of the fusion prediction module to achieve cross-modal attention weighted fusion is as follows: first, the speech feature signal, behavior feature signal and clinical feature signal are input into three independent fully connected layers for latent space projection alignment, and the features of different dimensions are mapped to the common latent space of unified dimension. The projection process adopts the orthogonal weight initialization strategy to ensure the orthogonality of features; then the cross-modal attention weighted operation is performed, and the query vector key vector and value vector are constructed for the three groups of modal combinations of speech and behavior, speech and clinical, and behavior and clinical, respectively. The similarity weight matrix between different modal features is calculated, and the similarity is converted into Attention distribution weights are used to dynamically adjust the contribution ratio of each modal feature in the fusion process; the weighted cross-modal interaction features are concatenated with the original modal features to form a joint feature vector, which is input into the gated recurrent unit network for temporal dependency modeling. Reset gates and update gates are set inside the network to control the flow of information. Layer normalization operations are introduced to accelerate convergence and residual connections are added to retain the underlying features, and finally a fused feature vector is output; this vector is input into the support vector machine classifier based on the radial basis kernel function, and the optimal classification hyperplane is solved through the sequential minimum optimization algorithm, and a continuous probability value in the range of zero to one is output as the initial risk prediction signal. The dynamic update module starts the incremental learning process when it detects a new data collection event: the feature extraction module performs the same processing process on the newly collected speech behavior and clinical data as the historical data to generate incremental speech features, incremental behavior features and incremental clinical features; these incremental features are projected into the latent space of the fusion prediction module, and then the cross-modal attention weights are calculated together with the historical fusion feature vectors, and the input gated recurrent unit is spliced ​​to generate the incremental fusion feature vector; then the online sequence extreme learning machine update mechanism is started, and the set of historical fusion feature vectors is used as the initial training set to construct the hidden layer output matrix, the input weights are randomly fixed and the rectified linear unit with leakage parameters is used for activation function, and solves the initial output weights through generalized inverse matrix operations; when a new fused feature vector arrives, the recursive generalized inverse formula is used to iteratively update the output weights, and the iterative process uses the matrix decomposition algorithm to accelerate the operation; the synchronous sliding time window mechanism maintains a fixed-capacity feature queue. When a new feature is added to the queue and the capacity limit is exceeded, the earliest historical feature is automatically eliminated, and the queue update triggers the local reconstruction of the hidden layer matrix; finally, the updated output weights are fed back to the support vector machine classifier of the fusion prediction module, and the hot start mechanism is used to adjust the kernel function parameters and dynamically scale the kernel function width coefficient according to the incremental data distribution. At the same time, the decision function bias term is corrected according to the proportion of new data categories.

[0028] Specifically, the process of converting the updated risk prediction signal into a visual decision support element by the visualization output module includes: the dynamic risk trajectory diagram generation stage, firstly sorting the historical and updated risk prediction signal sequences by timestamp, and using the non-uniform cubic spline interpolation algorithm to construct a continuous and smooth risk probability curve. The interpolation function satisfies the conditions that the function value is equal and the first-order and second-order derivatives are continuous in the interval of adjacent time points, and the spline coefficients are determined by solving the three-moment equations; marking clinical event points such as the doctor's intervention time and the drug adjustment time on the time axis and adding event icons above the curve, drawing horizontal reference lines to mark the low, medium and high risk threshold intervals; calculating the confidence interval of the prediction results at each time point based on the self-service sampling method, and drawing a circle around the curve. Generate a semi-transparent ribbon area; in the feature contribution heat map generation stage, the contribution of each modal feature to the current prediction result is quantified through the Shapley addition and interpretation algorithm. During the calculation, an eigenvalue perturbation matrix is ​​constructed to simulate the feature missing scenario, and the differences in the prediction output under different disturbance states are compared and the contribution value is attributed; the calculation results are mapped to the modal feature two-dimensional matrix, and the contribution intensity is represented by a red-green gradient color system, and the red area indicates the high-contribution feature; the feature contribution threshold is set to automatically trigger the associated playback function. When the contribution value of a specific acoustic feature exceeds the threshold, the original voice clip is associated and replayed and the waveform is superimposed. When the behavioral feature has a high contribution, the original signal curve of the sensor in the corresponding period is played back. When the clinical feature has a high contribution, a historical score comparison table pops up. The system deployment adopts a cloud-edge collaborative architecture: the multi-source data acquisition module runs on the user's smart terminal, voice acquisition uses the terminal microphone and real-time noise reduction algorithm, behavioral data acquisition calls the accelerometer, gyroscope and screen touch sensor, and clinical assessment data is entered through an encrypted form on the doctor's side; the feature extraction module is deployed on the edge computing node, and the voice acoustic feature extraction runs the optimized fast Fourier transform algorithm on the terminal digital signal processor, the semantic feature analysis is offloaded to the edge graphics processor server, the behavioral feature processing uses the hidden Markov model acceleration unit built into the sensor coprocessor, and the clinical feature generation is completed in the cloud central processing unit cluster; the fusion prediction module and the dynamic update module are deployed in the cloud high-performance server cluster, and resource elastic scheduling is achieved through containerization technology; the visualization output module is implemented based on web interaction technology, and a vector graphics library is used to render dynamic risk trajectories and heat maps; a quantum key distribution secure channel is established between the edge device and the cloud, and the feature signal is encrypted and encapsulated into a binary protocol data packet using the national secret algorithm before transmission, and the transport layer implements end-to-end integrity verification and anti-replay attack protection. All data flows between modules strictly follow timing logic, forming a closed-loop processing link from data acquisition to visualization output, ensuring the real-time and dynamic adaptability of schizophrenia risk prediction.

[0029] The present invention provides a system for early detection and risk prediction of schizophrenia. It collects multimodal data of speech, behavior and clinical assessment in natural situations in real time through a mobile terminal. After sub-modal feature extraction, it uses a cross-modal attention mechanism to dynamically weight and fuse speech acoustic semantic features, behavioral spatiotemporal pattern features and clinical structured features to generate a fused feature vector that is input into a machine learning classifier to output a risk prediction. It triggers incremental learning based on new data to update the classifier weights online to achieve dynamic risk assessment, and converts the prediction results into a visual dynamic risk trajectory map and feature contribution heat map to assist clinical decision-making, thus realizing a closed loop from continuous monitoring to interpretable early warning.

[0030] Therefore, the problem of delayed early screening of schizophrenia and lack of dynamic risk assessment can be solved through the system for early detection and risk prediction of schizophrenia of the present invention.

[0031] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the present invention. Anyone skilled in the art may modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by one of ordinary skill in the art without departing from the spirit and technical principles disclosed herein are intended to be covered by the claims of the present invention.

Claims

1. A system for early detection and risk prediction of schizophrenia, characterized in that: include: A multi-source data acquisition module, which generates a multimodal raw signal including speech data, behavioral digital phenotype data, and clinical assessment data; A feature extraction module, wherein the feature extraction module performs parallel processing on the multimodal original signal: extracting acoustic and semantic features of speech data to form a speech feature signal, extracting spatiotemporal pattern features of behavioral digital phenotype data to form a behavioral feature signal, and extracting structured indicators of clinical assessment data to form a clinical feature signal; A fusion prediction module, which fuses speech feature signals, behavioral feature signals, and clinical feature signals through a cross-modal attention weighting mechanism to generate a fusion feature vector, and outputs an initial risk prediction signal based on the vector through a machine learning classifier; A dynamic update module receives incremental feature signals triggered by newly collected user data, generates updated risk prediction signals through incremental learning, and feeds them back to the fusion prediction module to optimize the classifier weights; A visualization output module converts the updated risk prediction signal into a dynamic risk trajectory map and an interpretable feature contribution heat map.

2. A system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: The multi-source data acquisition module collects voice data by acquiring the original audio stream of natural conversation or specified text reading in real time through the mobile terminal microphone. The collection of behavioral digital phenotype data relies on the built-in sensors of the smart terminal to continuously monitor the user's movement trajectory, screen operation frequency and environmental interaction time series. The collection of clinical assessment data uses a structured electronic form to record the doctor's standardized scores of positive symptoms, negative symptoms and cognitive functions. The three types of data are encrypted and compressed and encapsulated into multimodal original signals with a unified timestamp.

3. The system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: When extracting speech feature signals, the feature extraction module first analyzes the fundamental frequency contour, formant distribution, and nonlinear dynamic parameters using a pre-trained acoustic model to generate an acoustic feature vector. Simultaneously, a deep semantic encoder is used to analyze the text coherence and logical structure abnormality to generate a semantic feature vector. The two are then concatenated and processed for dimensionality reduction to form a speech feature signal. In behavioral feature signal extraction, a time-frequency transformation algorithm is used to separate walking rhythm, gesture stability, and circadian activity intensity characteristics from the raw data of the accelerometer and gyroscope, and a Markov chain model is constructed to quantify the state transition probability; The clinical feature signals are then mapped into a continuous vector space through an embedding layer to map the discrete scale scores.

4. The system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: The cross-modal attention weighted mechanism of the fusion prediction module is specifically as follows: the speech feature signals, behavioral feature signals and clinical feature signals are projected into the common latent space respectively, the cross-attention weight matrix between each pair of modalities is calculated, and the cross-modal interaction features are generated according to the dynamic weighted summation of the weights. After being spliced ​​with the independent features of each modality, the cross-modal interaction features are input into the gated recurrent unit for time series modeling. The final output fusion feature vector is connected to a machine learning classifier composed of a support vector machine containing a radial basis function kernel.

5. The system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: The incremental learning process of the dynamic update module includes: after receiving the incremental feature signal triggered by the newly collected data, first use the same parameters of the feature extraction module to extract the incremental voice, behavior and clinical features, and then input them together with the historical fusion feature vector into the online sequence extreme learning machine, and update the classifier weights by iteratively solving the least squares solution of the hidden layer output weights. At the same time, a sliding time window mechanism is used to eliminate expired features to control the complexity of the model. The updated risk prediction signal is fed back to the fusion prediction module in real time to cover the old classifier.

6. The system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: When the visualization output module generates a dynamic risk trajectory diagram, the updated risk prediction signals at different time points are sorted by acquisition time, and a continuous risk probability curve is generated using cubic spline interpolation. The clinical intervention event nodes and symptom exacerbation threshold lines are marked on the curve. The feature contribution heat map quantifies the contribution of each modal feature to the current prediction result by calculating the shaplegi plus explanation value, maps it on the modality-feature two-dimensional matrix with a color gradient, and displays the original data fragments in association for verification.

7. The system for early detection and risk prediction of schizophrenia according to claim 3, characterized in that: The deep semantic encoder is constructed using a multi-head self-attention mechanism. Its input sequence is the word embedding vector after speech-to-text conversion. Contextual dependencies are captured by stacking bidirectional long short-term memory network layers. The output layer uses a convolutional neural network to extract local semantic patterns. The final semantic feature vector is the maximum pooling result of the hidden states of each network layer.

8. The system for early detection and risk prediction of schizophrenia according to claim 4, characterized in that: The gated recurrent unit introduces residual connections and layer normalization operations in the temporal modeling process. The weight matrices of the reset gate and update gate in its hidden state update function are dynamically initialized by cross-modal interaction features. The unit output is input into the support vector machine classifier after dropout regularization processing.

9. The system for early detection and risk prediction of schizophrenia according to claim 5, characterized in that: The hidden layer node activation function of the online sequential extreme learning machine adopts a rectified linear unit with a leakage parameter, its input layer is fully connected with the newly added feature signal and the historical fusion feature vector, the output layer error backpropagation only adjusts the weight of the last layer, and the weight matrix from the hidden layer to the output layer is iteratively updated by the Moore-Penrose generalized inverse matrix.

10. The system for early detection and risk prediction of schizophrenia according to claim 1, characterized in that: The system for early detection and risk prediction of schizophrenia is deployed in a cloud-edge collaborative architecture. The multi-source data acquisition module and the feature extraction module run on the user's intelligent terminal edge device. The fusion prediction module and the dynamic update module are deployed on the cloud server. The visualization output module is implemented through a web interactive interface. A secure channel based on quantum key distribution is used to transmit intermediate feature signals between the edge device and the cloud.

Citation Information

Cited By

  • Patient case analysis method and system

    CN121034510A

  • Intelligent home maintenance system

    CN121483590A

  • Intelligent mechanical control method and system based on deep learning

    CN121635059A