A depression risk assessment method and system based on big data

By constructing a unified time index sequence and multi-layer data processing technology, the modeling problem of the continuity and structural trajectory of multimodal behavioral data on the time axis was solved, realizing the stability and reliability of the mental health monitoring system, and enabling accurate assessment of depression risk over continuous time periods.

CN122337578APending Publication Date: 2026-07-03TIANJIN MEDICAL UNIV +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANJIN MEDICAL UNIV
Filing Date
2026-03-12
Publication Date
2026-07-03

Smart Images

  • Figure CN122337578A_ABST
    Figure CN122337578A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for assessing depression risk based on big data, relating to the field of big data behavioral analysis technology. The method includes constructing feature vectors for each time point, generating a vector set, setting a scale set using linear sampling and generating domain labels using a clustering algorithm, calculating a core score, generating difference sequences from the homology dimension and stacking them into a sequence matrix, extracting the first principal component using PCA to obtain a topological kernel vector and constructing a topological kernel matrix, obtaining residual perturbation vectors through singular value decomposition and performing discrete Fourier transform to obtain spectral values, calculating the split point index, and combining the stability score to obtain the depression risk. By constructing a manifold point cloud of behavioral vectors, the evolution trajectory of depression risk is accurately assessed. By extracting topological stability indicators and spectral clustering results, the detection accuracy and assessment reliability of persistent state anomalies are improved. The introduction of residual perturbation and frequency domain analysis mechanisms enhances the timeliness and accuracy of early warning of depression risk.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of big data behavioral analysis technology, and in particular to a method and system for assessing the risk of depression based on big data. Background Technology

[0002] Mental health monitoring methods are gradually evolving from traditional subjective questionnaires and clinical interviews to those based on big data and multimodal behavioral modeling. Research findings on the correlation between objective indicators such as heart rate variability, exercise intensity, sleep structure, vocal emotion, and application usage behavior are increasing. Through feature extraction within time windows, classification model training, and label regression, individual states are statically classified or segmented for prediction. Topological data analysis, graph signal processing, and spectral embedding techniques are receiving increasing attention in behavioral modeling, providing theoretical support for the geometric structure mining of temporal states.

[0003] However, existing technologies still have shortcomings. Traditional methods lack the ability to model the continuity, variability, and structural trajectory of multimodal behavioral data over time, making it difficult to depict the dynamic evolution of psychological states. Existing behavioral time-series processing structures used for mental health risk assessment, when actually deployed on the terminal or edge side, are limited by objective technical conditions such as inconsistent sampling frequencies, timestamp offsets, and unstable communication of multimodal sensor data. They usually use a single time scale or local statistical features for processing, making it difficult to maintain the consistency of cross-modal data structures during system operation. This leads to fluctuations or misjudgments in risk assessment results over continuous time periods, making it difficult to meet the technical requirements of stability and reliability for mental health monitoring systems. Summary of the Invention

[0004] In view of the aforementioned existing problems, the present invention is proposed.

[0005] Therefore, this invention provides a big data-based method and system for assessing the risk of depression. It addresses the shortcomings of traditional methods in modeling the continuity, variability, and structural trajectory of multimodal behavioral data over time, making it difficult to depict the dynamic evolution of psychological states. Existing behavioral time-series processing structures for mental health risk assessment, when deployed on the terminal or edge, are limited by objective technical conditions such as inconsistent sampling frequencies, timestamp offsets, and unstable communication of multimodal sensor data. They typically use a single time scale or local statistical features for processing, making it difficult to maintain consistency of cross-modal data structures during system operation. This results in fluctuations or misjudgments in risk assessment results over continuous time periods, failing to meet the technical requirements of stability and reliability for mental health monitoring systems.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides a method for assessing the risk of depression based on big data, comprising, Collect and preprocess multimodal data to generate discrete state sequences. Construct an adjacency graph from the discrete state sequences. Extract the adjacent points of each node based on the adjacency graph. Calculate the difference between the node and its adjacent points to generate a behavior vector field. Generate embedded coordinates by normalizing the Laplace matrix through random walk. Construct a set of time windows, extract and analyze the embedded coordinates of the time windows, generate local point clouds, calculate the Euclidean distance between any two local point clouds to obtain a distance matrix, set the number of filtering scales and extract effective distances, obtain a set of thresholds using the quantile method, perform edge filtering based on radius thresholds, construct a higher-order graph, extract complete subgraphs to obtain topological objects, use the standard PH algorithm to obtain a set of barcodes and calculate the lifetime length to obtain a set of window topological features, and convert the barcode distance into a stability score. Feature vectors are constructed for each time point to generate a vector set. A scale set is set using linear sampling and a clustering algorithm is used to generate domain labels. Core scores are calculated. Difference sequences are generated through homology dimensions and stacked into a sequence matrix. The first principal component is extracted using PCA to obtain the topological kernel vector and construct the topological kernel matrix. The residual perturbation vector is obtained through singular value decomposition and the spectrum value is obtained by performing discrete Fourier transform. The split point index is calculated and combined with the stability score to obtain the depression risk.

[0007] As a preferred embodiment of the big data-based depression risk assessment method of the present invention, the method comprises: calculating the difference between the computing node and its neighboring nodes to generate a behavioral vector field, and generating embedded coordinates through a random walk normalized Laplacian matrix, including: A data matrix is ​​constructed for the discrete state sequence. The Euclidean distance between any two rows of the data matrix is ​​calculated using the Euclidean distance formula. The Euclidean distances are sorted in ascending order. The time points of the two rows corresponding to the smallest Euclidean distance are selected and marked as adjacency relationships. An adjacency graph is constructed. All adjacent points of each node are extracted. The difference between each node and all its corresponding adjacent points is calculated to obtain the state difference. The state difference is normalized to generate a behavior evolution direction vector. The behavior evolution direction vector is sorted according to the time points to obtain the behavior vector field. The edge weights are defined as rows, and the state vectors are defined as columns. A similarity matrix is ​​constructed, and a graph Laplacian matrix is ​​constructed using the random walk normalized Laplacian matrix method. The graph Laplacian matrix is ​​then decomposed into eigenvalues ​​to obtain the feature components and generate the embedded coordinates.

[0008] As a preferred embodiment of the big data-based depression risk assessment method of the present invention, the step of obtaining a threshold set through the quantile method, performing edge filtering based on the radius threshold, and constructing a higher-order graph includes: Construct a set of time windows, extract the embedding coordinates of the popular point clouds of the analysis time windows in the set of time windows, and obtain a set of local point clouds; For any two local point clouds in the local point cloud set, calculate the Euclidean distance using the Euclidean distance formula and generate a distance matrix; The number of filtering scales is set based on empirical rules. Effective distances are extracted based on the distance matrix. The starting and ending radius values ​​are calculated using the quantile method. Linear interpolation is performed on the starting and ending radius values ​​to obtain a threshold set. For each radius threshold in the threshold set, the local point cloud in the distance matrix is ​​traversed to construct a high-order graph.

[0009] As a preferred embodiment of the big data-based depression risk assessment method of the present invention, the step of converting barcode distance into a stability score includes: Extract complete subgraphs from higher-order graphs, generate topological objects, combine the Euclidean norm values ​​of the behavior vector field, and use the standard PH algorithm to calculate the barcode set in each homology dimension, including 0-dimensional, 1-dimensional, and 2-dimensional, and calculate the barcode weights. The barcode weights of each homology dimension in each analysis time window are summed to obtain the stability contribution value of each window. The stability contribution values ​​of all homology dimensions are weighted and summed to obtain the topological stability index and the window topological feature set. The distance between the window topological feature set and the healthy topological feature set is calculated using the barcode distance formula and converted into a stability score using the natural exponential function.

[0010] As a preferred embodiment of the big data-based depression risk assessment method described in this invention, the step of using a linear sampling method to set a scale set and using a clustering algorithm to generate domain labels, and calculating a core score, includes: A feature vector is constructed for each time point to obtain a vector set. A linear sampling method is used to set the scale parameter. For each scale parameter, a clustering algorithm is used to cluster the vector set to obtain the domain label. For each time point, iterate through each scale parameter in the scale set, select the domain label that appears most frequently, set it as the structural domain label, and calculate the core score.

[0011] As a preferred embodiment of the big data-based depression risk assessment method of the present invention, the calculation of the split point index, combined with the stability score to obtain the depression risk, includes: Based on the number of filtering scales, the homology dimensions under each filtering scale are arranged vertically to obtain curve vectors. First-order difference is performed on the curve vectors to obtain sequence matrices. Principal component analysis is performed on the matrix to extract the first principal component, which is set as the topological kernel vector. A topological kernel matrix is ​​constructed, and singular value decomposition is performed on the topological kernel matrix to calculate the residual perturbation vector. The discrete Fourier transform of the residual perturbation vector is used to obtain the spectral value, the split point index is calculated, and the overall structural stability is calculated in combination with the stability score. The overall structural stability is normalized to obtain the stability trajectory, and the risk of depression is calculated.

[0012] As a preferred embodiment of the big data-based depression risk assessment method of the present invention, the step of collecting multimodal data and preprocessing it to generate a discrete state sequence includes: Multimodal data is collected through API interfaces, including heart rate, steps, average acceleration, total application usage time, sleep assignment tags, total screen on and off time, call logs, voice diaries, text, environmental data, genetic data, image feature vectors, age, and gender. Discrete state sequences are generated from multimodal data through preprocessing, including constructing a time index sequence, horizontally stacking the preprocessed multimodal data to obtain a state vector, normalizing it, and incorporating it into the time index sequence to generate a discrete state sequence.

[0013] Secondly, the present invention provides a big data-based depression risk assessment system, comprising, The multimodal acquisition and preprocessing module is used to collect multimodal data through the API interface and generate preprocessed multimodal data using sliding window analysis and forward filling. The discrete state sequence construction module is used to construct time-indexed sequences by horizontally stacking and normalizing preprocessed multimodal data to generate discrete state sequences. The graph embedding and behavior vector field generation module is used to construct a data matrix based on discrete state sequences, generate an adjacency graph, a similarity matrix, embedding coordinates, and a behavior vector field, and obtain a popular point cloud. The Time Window Topology Analysis and Stability Scoring module is used to construct a set of time windows and generate topology objects, and to calculate the barcode set, barcode weight, topology stability index and stability score; The multi-scale structural domain identification and depression risk calculation module is used to construct feature vectors and generate domain labels, and calculate core degree score, topological kernel vector, split point index and comprehensive structural stability to obtain depression risk.

[0014] Thirdly, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the big data-based depression risk assessment method described in the first aspect of the present invention.

[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the big data-based depression risk assessment method described in the first aspect of the present invention.

[0016] The beneficial effects of this invention are as follows: Addressing the issues of inconsistent sampling frequencies, timestamp offsets, and unstable communication in multimodal sensor data within mental health monitoring systems, this invention constructs a unified time index sequence and preprocesses multi-source monitoring data to achieve alignment and fusion of data collected from different devices on a time scale. This eliminates risk assessment jitter caused by asynchronous multi-source data and improves the structural consistency of monitoring data in continuous time series. Furthermore, by performing structured modeling and stability analysis on multimodal behavioral data, the system can stably depict the changing trends of individual behavioral states within continuous time windows, reducing the impact of short-term noise or occasional anomalies on assessment results. Simultaneously, by analyzing the persistence of behavioral structural changes, it can identify small-amplitude but long-lasting abnormal behavioral changes in the early stages of depression, thereby improving the mental health monitoring system's ability to identify trends in depression risk and its early warning capabilities, making risk assessment results more stable and reliable in long-term continuous monitoring environments. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of the big data-based depression risk assessment method in Example 1.

[0019] Figure 2 This is a schematic diagram of the depression risk assessment system based on big data in Example 1. Detailed Implementation

[0020] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0021] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0022] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0023] Example 1, referring to Figure 1 and Figure 2 This is the first embodiment of the present invention, which provides a method for assessing the risk of depression based on big data, including the following steps: S1. Collect multimodal data and preprocess it to generate a discrete state sequence. Construct an adjacency graph for the discrete state sequence. Extract the adjacent points of each node based on the adjacency graph. Calculate the difference between the node and its adjacent points to generate a behavior vector field. Generate embedded coordinates by normalizing the Laplace matrix through random walk. Specifically, multimodal data is collected and preprocessed to generate discrete state sequences, including: In mental health monitoring systems, different monitoring devices typically have different sampling periods. For example, wearable devices collect heart rate data at a frequency of seconds, while mobile terminals record behavioral activities at a frequency of minutes. The sampling period for environmental sensing devices may range from several minutes to tens of minutes. Therefore, there is a significant inconsistency between different data sources in the time dimension. If the data is directly fused, it is easy to cause misalignment of the data on the time axis, resulting in significant fluctuations in the behavioral state assessment results over a continuous period of time. This invention constructs a unified time index sequence and performs time alignment processing on multimodal monitoring data, mapping data from different devices to a unified time scale. At the same time, it combines sliding window missing compensation and forward filling denoising processing to suppress packet loss and random noise generated during communication, enabling multi-source monitoring data to form a continuous and structurally consistent discrete state sequence, thereby improving the stability of the mental health monitoring system during continuous operation. Multimodal data is collected through API interfaces, and discrete state sequences are generated through preprocessing. The multimodal data refers to data with clear physical or behavioral meaning collected by wearable devices, mobile terminals, and environmental sensing devices, including heart rate (mean, variance, and HRV), steps, mean acceleration, total application usage time, sleep assignment tags (one-hot encoding), total screen on and off time, call logs (number of calls and duration), voice diaries (acoustic features, including fundamental frequency extraction via YIN algorithm, energy obtained by summation of squares, MFCC extraction via Mel-frequency cepstral coefficients, and emotion score extraction via random forest), text (text emotion score obtained via support vector machine, extraction of negative word proportion), environmental data, genetic data, image feature vectors, age, and gender (one-hot encoding). The genetic data includes single nucleotide polymorphism (SNP) risk sets (each SNP is encoded as 0 / 1 / 2 (allele count)), epigenetic methylation levels (each CpG site is normalized), candidate gene expression risk scores, family genetic risk index, PRS polygenic risk score, and inflammation-related gene markers (normalized). The environmental data includes temperature, light intensity, noise, outdoor air quality index, PM2.5 and PM10 concentrations, humidity, air pressure, ultraviolet index, average community noise level, geographical location and activity area density, and social environment index. The image feature vector includes facial expression images, everyday images, and medical images collected through an API interface. Facial expression probabilities (e.g., happiness and sadness) are extracted using a CNN on the facial expression images. AU facial action units are obtained through an AU detection model, and the feature vector is calculated using Eye Aspect Ratio (EAR) + Dark Circle. The system detects dark circles under the eyes, obtains facial fatigue index (e.g., color difference of eye bags), extracts scene brightness and calculates tone entropy from everyday scene images, extracts the number of people in social situations using the YOLO model, extracts brain structure volume (e.g., amygdala) and cortical thickness from structural MRI in medical images using FreeSurfer, and calculates local consistency, low-frequency amplitude, and functional connectivity strength of key brain networks (e.g., affective and executive control networks) from resting-state functional MRI data in medical images using Kendall's concordance coefficient method, and extracts the functional connectivity strength of key brain networks (e.g., affective and executive control networks) based on seed point correlation analysis or independent component analysis. The system first normalizes expression probability, facial action units (AU), facial fatigue index, scene brightness, tone entropy, number of people in social situations, brain structure volume, cortical thickness, local consistency, low-frequency amplitude, and functional connectivity strength, then stacks them horizontally to obtain image feature vectors. The preprocessing includes filling missing values ​​in the multimodal data using sliding window analysis to address the problem of inconsistent sampling frequencies of multi-source sensors, and denoising the data loss and noise generated during wireless communication by forward padding to obtain preprocessed multimodal data. The discrete state sequence includes constructing a time index sequence, such as t1, t2...tn (t is a time point, t1 is the first time point, and n is the number of time points), stacking the preprocessed multimodal data horizontally to obtain the state vector at time t1, normalizing it, and incorporating it into the time index sequence to generate the discrete state sequence.

[0024] By aggregating multimodal data in real time through API interfaces and using a sliding window approach to fill in missing values ​​and perform forward imputation and denoising, the unified temporalization and data standardization of heterogeneous behaviors and physiological information are achieved. This significantly enhances the structural usability and modeling continuity of the data. The horizontal stacking of multidimensional features such as speech fundamental frequency, text emotion, and application behavior constructs a normalized state sequence, laying a rich and consistent foundation for subsequent structural modeling. This solves the problem of inconsistent sampling frequencies and heterogeneous distribution of multi-source signals, making it difficult to model, and improves the perceptibility of psychological state evolution under micro-perturbations.

[0025] Furthermore, the differences between nodes and their neighbors are calculated to generate a behavior vector field. Embedded coordinates are generated by normalizing the Laplacian matrix through random walks, including: Construct a data matrix for the discrete state sequence, where the rows are state vectors and the columns are the eigenvalues ​​of the state vectors; Calculate the Euclidean distance between any two rows of the data matrix using the Euclidean distance formula, sort the Euclidean distances in ascending order, and select the time points of the two rows corresponding to the smallest Euclidean distance, marking them as adjacency relationships. Define time points as nodes. If there is an adjacency relationship between nodes, then set there to be an edge between nodes. Calculate the weight between nodes as the natural exponential function value of the negative Euclidean distance to construct an adjacency graph. Based on the adjacency graph, all adjacent nodes of each node are extracted, and the difference between each node and all its corresponding adjacent nodes is calculated to obtain the state difference. This difference is then normalized to generate a behavior evolution direction vector. This vector is then sorted according to time points to obtain the behavior vector field, as shown in the formula: , in, The evolution direction vector, Let be the state vector in the discrete state sequence. For the first At a certain point in time, For the first The set of neighbors at each time point is extracted from the adjacency graph. For the first The time point and the The similarity weights at each time point are called edge weights. For the first At a certain point in time, For state differences; By calculating the state difference between adjacent time points in a discrete state sequence and constructing a behavior vector field, the changing trend of multimodal behavior data can be represented as a directional evolutionary trajectory. This processing method can effectively suppress the influence of instantaneous outliers from a single sensor on the overall behavior state judgment, enabling the system to more stably depict the changing trend of individual behavior over a continuous time period. By introducing an adjacency graph and similarity weight mechanism, the structural bias caused by noise from individual data points can be further reduced, improving the stability of behavior trajectory modeling. Edge weights are defined as rows, and state vectors are defined as columns to construct a similarity matrix. A graph Laplacian matrix is ​​then constructed using the random walk normalized Laplacian matrix method. Eigenvalue decomposition is performed on the graph Laplacian matrix to obtain feature components, generating embedded coordinates (which are then standardized). These embedded coordinates are sorted by time point to obtain the popular point cloud, as shown in the formula: , in, For embedding coordinates, For transpose, For the first Each point in time in dimension The coordinate values ​​are the characteristic components.

[0026] Based on the constructed data matrix, adjacency relationships between time points are established using the minimum Euclidean distance as the criterion, effectively capturing the similarity features between adjacent behavioral states. The subsequently constructed adjacency graph and behavioral vector field not only preserve the continuity and directionality of state evolution, but also suppress the interference of non-structural mutations on the system through similarity weighting. By using the random walk normalized Laplace matrix, a compressed mapping of the complex state space is achieved, enabling high-dimensional behavioral paths to represent manifold structures in low-dimensional coordinate form, which greatly enhances the interpretability of subsequent topological stability analysis and the accuracy of dynamic trajectory analysis.

[0027] S2. Construct a set of time windows, extract the embedded coordinates of the analysis time windows, generate local point clouds, calculate the Euclidean distance between any two local point clouds, obtain the distance matrix, set the number of filtering scales and extract the effective distances, obtain the threshold set through the quantile method, perform edge filtering based on the radius threshold, construct a higher-order graph, extract the complete subgraph, obtain the topological objects, use the standard PH algorithm to obtain the barcode set and calculate the lifetime length, obtain the window topological feature set, and convert the barcode distance into a stability score. Specifically, a threshold set is obtained through the quantile method, edges are filtered based on radius thresholds, and a higher-order graph is constructed, including: The initial analysis window length is set based on the rule of thumb. The time point t1 and the initial analysis window length are added together to obtain the analysis time point t1. The time interval from time point 1 to analysis time point t1 is the analysis time window. The same operation is performed on the remaining time points. The obtained analysis time windows are arranged horizontally to obtain the time window set. The embedding coordinates of the popular point clouds of the analysis time windows are extracted from the time window set to obtain local point clouds. The local point clouds are then arranged horizontally to obtain the local point cloud set of each analysis time window (the embedding coordinates in the set are standardized). The local point cloud is a set of behavioral state embedding points obtained by embedding multimodal monitoring data within a fixed analysis time window during the depression risk assessment process. It is used to characterize the distribution characteristics of behavior and physiological state in a low-dimensional structural space within the time window. For any two local point clouds in the set of local point clouds, calculate the Euclidean distance using the Euclidean distance formula and generate a distance matrix, where the rows and columns are the p-th local point cloud and the k-th local point cloud, and the matrix elements are the Euclidean distances; The number of filtering scales is determined based on empirical rules. Effective distances are extracted based on the distance matrix. The starting and ending radius values ​​are calculated using the quantile method. Linear interpolation of the starting and ending radius values ​​yields the radius threshold for each filtering scale. These thresholds are then arranged horizontally to obtain a set. The formula is as follows: , , , in, For effective distance, For the first The first analysis time window Local point cloud and the first Euclidean distance of a local point cloud For the first Number of local point clouds in each analysis time window This is the initial radius value. The termination radius value, This indicates that the Euclidean distances in the distance matrix are sorted in ascending order, and the selection position is... The value at the quantile, and These are the minimum quantile and maximum quantile parameters, which are empirical values. For each radius threshold in the threshold set, traverse the local point cloud in the distance matrix, filter the local point clouds in the distance matrix whose Euclidean distance is less than or equal to the radius threshold, then assume that there is an edge between the local point clouds, arrange all the edges in the distance matrix that meet the filtering conditions horizontally to obtain the edge set, define the local point cloud set as the point set, and obtain the higher-order graph. The higher-order graph is a connection structure built based on the distance relationship between behavioral state embedding points within a time window. It is used to describe the association pattern of multiple behavioral states at the same time scale. Its purpose is to assist in the calculation of the stability of the behavioral structure during the depression risk assessment process. By setting an initial analysis window and extracting local point clouds in segments, the structural heterogeneity under the background of temporal evolution is effectively characterized. The Euclidean distance matrix combined with the quantile method is used to set the filtering scale, and a multi-scale radius threshold set is introduced to enable the structural sensitivity adjustment to have dynamic adaptability. On this basis, a high-order graph is constructed between point clouds, which fully explores the dense connectivity patterns and geometric topological characteristics of the group behavior structure. This helps to reveal potential ring-shaped, clustered, or cavity behavioral blocks, providing accurate boundaries and geometric support for subsequent complex construction and barcode extraction, and effectively improving the system's robust modeling ability for the evolution of potential psychological state structures.

[0028] Furthermore, the barcode distance is converted into a stability score, including: Extract the complete subgraph from the higher-order graph to represent any n local point clouds, where there is an edge connecting every pair of them. Arrange the local point clouds in the complete subgraph horizontally to obtain a simplex, and generate a set of simplexes for the analysis time window. The complete subgraph is a structural unit of highly correlated behavioral states obtained by screening in a high-order graph. It means that multiple behavioral states within the same analysis time window all meet the preset correlation conditions, and is used to identify state combinations with stable characteristics in the process of depression risk assessment. Define the set of simplexes, edges, and points as complexes, and generate topological objects using the following formula: , in, For the first A complex analysis time window, For topology objects, For the first The radius threshold for each filtering scale. The number of filter scales; The topology object is a structural model constructed based on the embedding points of behavior states within a time window and their correlations, used to calculate and analyze the structural stability of multimodal monitoring data at different scales; In multimodal behavioral data, different types of data features often have different scales of change. For example, the magnitude of change in motion data is usually greater than that of speech emotion features or text emotion features. If only a single scale is relied upon for analysis, it is easy to overlook some potential behavioral structures. By constructing a multi-scale filtering radius and building a high-order graph structure at different scales, the system can simultaneously analyze the correlation between behavioral states at multiple spatial scales, thereby improving the system's ability to recognize complex behavioral patterns. The multi-scale topology structure can effectively reduce the interference of local abnormal data on the overall structure judgment, making the assessment of behavioral structure stability more reliable. Based on topological objects and vector norms (the Euclidean norm values ​​of behavioral vector fields, set as weights), the standard PH algorithm is used to calculate the barcode set in each homology dimension; The barcode set is a set of scale intervals corresponding to the generation and disappearance of behavioral structural features formed by multimodal monitoring data within a time window under different filtering scales. It is used to describe the persistence and stability of behavior and physiological state at the structural level. It is an intermediate calculation result in the depression risk assessment system and does not involve direct judgment of individual psychological state or behavioral semantics. The homology dimension includes 0-dimensional (connected components), 1-dimensional (ring structure), and 2-dimensional (cavity structure). For each interval of the barcode set in each homology dimension (divided by setting the radius threshold as the dividing point), the lifetime length is obtained by calculating the start and end values ​​of the interval; The lifetime length is the length of the scale interval corresponding to the generation and disappearance of behavioral structural features under different filtering scales. It is used to quantify the stability of the behavioral state structure within the time window. Its value reflects the characteristics of the data structure, rather than the duration of the individual's psychological state. Healthy population samples are collected via API interface. Statistical analysis is used to set a minimum lifespan threshold and a set of healthy topological features. If the lifespan is less than or equal to the minimum lifespan threshold, the barcode weight is set to 0; otherwise, the barcode weight is calculated using the following formula: , in, For the first The first homology dimension The barcode weight of each barcode. For the first Importance coefficients of each homology dimension For the first The first homology dimension The lifespan of a barcode. For the first The maximum lifetime length of each homogeneous dimension is obtained by sorting the lifetime lengths in descending order. It is a very small positive number. For the first Minimum lifetime threshold for each homogeneous dimension; Barcode weights are used to characterize the stability of behavioral structures at different scales. When multimodal monitoring data is affected by short-term noise or anomalous sampling, the corresponding topological structure usually exists only within a very small scale, so its barcode lifetime is short and the corresponding weight is small. However, when the individual behavioral state shows a stable trend, its topological structure persists within a larger scale and the corresponding weight is large. By introducing a barcode weight mechanism, the impact of random noise on the structural analysis results can be effectively reduced, thereby improving the reliability of behavioral structure stability assessment. The barcode weights of each homology dimension in each analysis time window are summed to obtain the stability contribution value of each window. The stability contribution values ​​of all homology dimensions are then weighted and summed to obtain the topological stability index. Topological stability metrics are used to characterize the overall structural stability of multimodal behavioral data within a time window. When an individual's behavioral state remains stable, its embedded point cloud structure exhibits a persistent connected structure or ring structure in the multi-scale topological space, thus resulting in a high corresponding topological stability metric. When behavioral patterns undergo abnormal changes, such as a sudden decrease in activity levels or a change in sleep structure, the topological structure will break down or be reconstructed, and the corresponding topological stability metric will decrease significantly. By introducing topological stability metrics, the trend of behavioral structural changes can be stably characterized in continuous time series, thereby improving the system's ability to identify changes in the risk of depression. By weighting the barcode lifetime length to form barcode weights and further constructing a topological stability index, the stability of behavioral structures can be quantitatively analyzed in multi-scale topological spaces. This not only reduces the impact of short-term noise and sensor sampling errors on the evaluation results, but also captures the gradual change trend of behavioral structures in continuous time series, thereby improving the ability of mental health monitoring systems to identify early behavioral abnormalities in depression. The topological stability index and the barcode set are arranged horizontally to obtain the window topological feature set. The distance between the window topological feature set and the healthy topological feature set is calculated using the barcode distance formula, and then converted into a stability score using the natural exponential function. The stability score is used to quantify the continuity and consistency of multimodal monitoring data in time structure. When multi-source monitoring data is affected by sensor error, environmental noise or communication delay, the behavioral structure within a local time window may fluctuate briefly. By comparing the degree of difference between the current time window topology and the baseline topology of healthy people, short-term anomalies caused by noise can be effectively identified and smoothed through the stability score, thereby reducing the probability of misjudgment in the continuous evaluation process of the system. By constructing complete subgraphs in higher-order graphs and generating simplex sets and topological objects, the system captures the stability patterns of behavioral trajectories in spatiotemporal evolution from different scales and multidimensional homology structures. It achieves deep modeling of point cloud geometry at the topological level, effectively identifying latent patterns and stable configurations in behavioral structures. By applying persistent homology methods across multiple homology dimensions, it accurately characterizes the generation and disappearance processes of behavioral states at different filtering scales, significantly improving sensitivity to structural perturbations and risk precursors. Combined with the topological baseline map of healthy individuals, barcode weights, as an adjustable quantitative tool with bias control capabilities, achieve a balance between individualized modeling and group control in the assessment process, exhibiting good generalization and clinical adaptability. Jointly comparing stability indicators with barcode sets and converting them into stability scores through barcode distance calculation provides continuous risk expression while maintaining topological integrity. This effectively avoids the boundary uncertainties and threshold sensitivity faced by traditional methods based on discrete classification, providing a solid geometric foundation and quantitative tool system for the temporal stability modeling of depression.

[0029] S3. Construct feature vectors for each time point, generate vector sets, use linear sampling to set scale sets and use clustering algorithms to generate domain labels, calculate core score, generate difference sequences for homology dimension and stack them into sequence matrix, extract the first principal component through PCA, obtain topological kernel vector and construct topological kernel matrix, obtain residual perturbation vector through singular value decomposition, and perform discrete Fourier transform to obtain spectrum value, calculate split point index, and combine stability score to obtain depression risk; Specifically, a scale set is defined using linear sampling and a clustering algorithm is used to generate domain labels. Core scores are then calculated, including: For each time point, a feature vector is constructed, including embedding coordinates, homology dimension, topological stability index, vector norm (Euclidean norm value of behavioral vector field), and stability score. The feature vectors are then arranged horizontally to obtain standardized vectors. The standardized vectors are then sorted by time point to obtain a vector set. The feature vector is a time-point behavioral state feature vector constructed for depression risk assessment. It is formed by combining physiological monitoring parameters, behavioral activity parameters and structural stability parameters corresponding to the same time point, and is used to characterize the comprehensive characteristics of an individual's behavior and physiological state at that time point. A linear sampling method is used to define the scale set, which includes K scale parameters; For each scale parameter, a clustering algorithm (such as Gaussian Mixture Model (GMM) or spectral clustering) is used to cluster the vector set (the number of clusters is set by the scale parameter; for example, the larger the scale, the fewer the number of clusters) to obtain the domain labels. For each time point, iterate through each scale parameter in the scale set, select the domain label that appears most frequently, set it as the structural domain label, and calculate the core score using the following formula: , in, Assess the core score. For scale The most frequently occurring domain tags include 1 for active socializing or exercise, 2 for quiet rest or reading, and 3 for low mood or slow behavior. The structural domain labels are behavioral state structural interval identifiers obtained through multi-scale clustering methods, used to distinguish the structural distribution types of behavior and physiological states in different time periods during the depression risk assessment process; The core score is used to represent the degree of consistency in the behavioral structure segmentation at different scales at the same time point. Its value reflects the stability level of the behavioral state at that time point under continuous sampling conditions and is used for abnormal state identification in depression risk assessment. By constructing a feature vector set composed of embedded coordinates, homology dimension, and topological stability index, and performing multi-scale sampling and clustering analysis, the system can effectively uncover the structural aggregation and state distribution patterns of behavioral trajectories at different granular scales. This overcomes the traditional limitations of clustering at a single time scale. The introduction of a linear sampling strategy allows the system to adaptively discover potential structural domains over a wide time interval, thereby achieving spatial nesting of dynamic behavioral features. The domain labels generated by clustering are transformed into structural domain labels by statistically analyzing their frequency of occurrence, effectively avoiding the randomness and discreteness in the clustering results. This makes the structural domain identification process more stable and robust. By calculating the clustering consistency at all scales for each time point, a core score is formed. This score has a dual interpretive power in expressing the stability of individual behavioral states. On the one hand, it reflects the aggregation centrality of individuals in multi-scale space; on the other hand, it can serve as a warning signal for abnormal behavior deviating from the structural backbone. By embedding dynamic behavioral semantics while maintaining multi-scale structural consistency, the system effectively bridges the semantic gap between low-dimensional clustering labels and high-dimensional behavioral representations, providing a clear structural prior for subsequent identification of emotional fluctuation trends. This enhances the system's ability to identify and provide early warnings of state transition edge regions.

[0030] Furthermore, the split point index is calculated and combined with the stability score to obtain the risk of depression, including: Based on the number of filtering scales, the homology dimensions under each filtering scale are arranged vertically to obtain curve vectors. The first-order difference of the curve vectors is then performed to obtain difference vectors, which are arranged according to the filtering scales to generate a difference sequence. Extract 0-dimensional (connected components), 1-dimensional (loop structure), and 2-dimensional (cavity structure) difference sequences from the difference sequences respectively. Stack the difference sequences to obtain a sequence matrix. Perform principal component analysis on the matrix to extract the first principal component, set it as the topological kernel vector, and construct the topological kernel matrix. The formula is as follows: , in, For the first Topological kernel matrix for each analysis time window, For the first Topological kernel vectors for each analysis time window, For the first Topological kernel vectors for each analysis time window; The topological kernel vector is the principal representation vector extracted from multi-scale behavioral structural features through principal component analysis, which is used to reflect the main changing trend of behavioral state in the time structure during the depression risk assessment process. Perform singular value decomposition on the topological kernel matrix and calculate the residual perturbation vector using the following formula: , , in, For the first Topological skeleton matrix for each analysis time window, For the first Transpose of the right singular matrix of each analysis time window Represents the first column vector. For the first The left singular matrix of each analysis time window, For the first One row of the topological skeleton matrix for each analysis time window. For the first The residual perturbation vector for each analysis time window; The discrete Fourier transform is applied to the residual perturbation vector to obtain the spectral values, and the split point exponent is calculated using the following formula: , , , in, For the first High-frequency energy within an analysis time window, For the first Spectral values ​​for each analysis time window, For frequency, A set of non-zero frequencies This refers to the change in high-frequency energy. For the first High-frequency energy within an analysis time window, For column index, The mean of the core score. For the first Each analysis time window corresponds to a specific time point; The split point index is a structural perturbation index calculated based on the frequency domain characteristics of the residual perturbation vector. It is used to quantify the degree of abrupt changes in the temporal structure of behavioral and physiological data, so as to help identify structural instability caused by abnormal behavior or physiological fluctuations during the depression risk assessment process. By integrating topological stability indicators, structural consistency indicators, and breakpoint indices, a continuous risk evolution trajectory can be constructed, enabling the analysis of behavioral state change trends on a continuous time scale. This improves the system's ability to identify early behavioral abnormalities in depression. When individual behavioral patterns show persistent and subtle changes, the system can identify risk change trends in advance, thus providing earlier warning information for mental health intervention. The overall structural stability is calculated based on the split point index and stability score. The stability trajectory is then obtained by normalizing the overall structural stability. The risk of depression is then calculated using the following formula: , , in, To ensure overall structural stability, , as well as These are the weighting coefficients. It is a topological stability index. For the risk of depression, The maximum change in stability is obtained based on statistical analysis. The comprehensive structural stability is a composite index calculated based on the behavioral structural characteristics formed by multimodal monitoring data within a time window, used to quantify the continuity and consistency of behavioral and physiological data at the time structure level. By performing differential analysis on the homology structure under multi-scale filtering and extracting its principal components, a topological kernel vector reflecting the trend of behavioral changes is constructed. Furthermore, singular value decomposition is used to capture its residual perturbation information, which is then converted into spectral energy through Fourier transform. This enhances the system's response to structural perturbations and overcomes the problem of difficulty in explaining local fluctuations in traditional time series modeling. The introduction of the split point index, as a coupling quantity between high-frequency energy changes and behavioral structural consistency, is highly sensitive in reflecting key structural abrupt events, particularly suitable for capturing periodic structural deconstruction or rapid emotional breakdown processes. The split point index, linked with core score, forms a split point that... The point control mechanism not only enhances the system's sensitivity to the transition from stable to unstable processes, but also provides a structural prior basis for setting risk thresholds. The system integrates topological stability scores, structural consistency indicators, and split-point perturbation characteristics to form a comprehensive structural stability trajectory. This trajectory is transformed into a depression risk assessment curve through normalization, which not only avoids the risk misjudgment problem caused by the bias of a single indicator, but also establishes the temporal continuity and structural dependence correlation of the risk trajectory, making the risk assessment results have significant behavioral interpretability and evolutionary traceability, thus providing high-resolution decision support for early warning and dynamic intervention strategies. To address practical technical issues in mental health monitoring systems, such as inconsistent sampling frequencies, timestamp offsets, and communication noise from multiple monitoring devices, this paper proposes a structured fusion processing method for multimodal monitoring data by constructing a unified time index sequence, modeling the behavioral vector field structure, and performing multi-scale topological stability analysis. This method not only eliminates inconsistencies in the time dimension of multi-source data but also improves the system's ability to identify subtle behavioral abnormalities in the early stages of depression, thereby significantly enhancing the stability and reliability of the mental health monitoring system in a continuous operating environment.

[0031] This embodiment also provides a big data-based depression risk assessment system, including: The multimodal acquisition and preprocessing module is used to collect multimodal data through the API interface and generate preprocessed multimodal data using sliding window analysis and forward filling. The discrete state sequence construction module is used to construct time-indexed sequences by horizontally stacking and normalizing preprocessed multimodal data to generate discrete state sequences. The graph embedding and behavior vector field generation module is used to construct a data matrix based on discrete state sequences, generate an adjacency graph, a similarity matrix, embedding coordinates, and a behavior vector field, and obtain a popular point cloud. The Time Window Topology Analysis and Stability Scoring module is used to construct a set of time windows and generate topology objects, and to calculate the barcode set, barcode weight, topology stability index and stability score; The multi-scale structural domain identification and depression risk calculation module is used to construct feature vectors and generate domain labels, and calculate core degree score, topological kernel vector, split point index and comprehensive structural stability to obtain depression risk.

[0032] This embodiment also provides a computer device applicable to the big data-based depression risk assessment method, comprising: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the big data-based depression risk assessment method proposed in the above embodiment.

[0033] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.

[0034] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the big data-based depression risk assessment method proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0035] In summary, this invention addresses the issues of inconsistent sampling frequencies, timestamp offsets, and unstable communication in multimodal sensor data within mental health monitoring systems. By constructing a unified time index sequence and preprocessing multi-source monitoring data, it achieves alignment and fusion of data collected from different devices on a time scale. This eliminates risk assessment jitter caused by asynchronous multi-source data and improves the structural consistency of monitoring data in continuous time series. Furthermore, by performing structured modeling and stability analysis on multimodal behavioral data, the system can stably depict the changing trends of individual behavioral states within continuous time windows, reducing the impact of short-term noise or occasional anomalies on assessment results. Simultaneously, by analyzing the persistence of behavioral structural changes, it can identify small-amplitude but long-lasting abnormal behavioral changes in the early stages of depression. This enhances the mental health monitoring system's ability to identify trends in depression risk and provides early warnings, making risk assessment results more stable and reliable in long-term continuous monitoring environments.

[0036] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A method for assessing the risk of depression based on big data, characterized in that: include, Multimodal data is collected through API interfaces, including heart rate, steps, average acceleration, total application usage time, sleep assignment tags, total screen on and off time, call logs, voice diaries, text, environmental data, genetic data, image feature vectors, age, and gender. The multimodal data is preprocessed to generate a discrete state sequence, including constructing a time index sequence, stacking the preprocessed multimodal data horizontally to obtain a state vector, normalizing it, incorporating it into the time index sequence to generate a discrete state sequence, constructing an adjacency graph for the discrete state sequence, extracting the adjacent points of each node, calculating the difference between the node and its adjacent points to generate a behavior vector field, and generating embedded coordinates by random walk normalizing the Laplace matrix. Construct a set of time windows, extract and analyze the embedded coordinates of the time windows, generate local point clouds, calculate the Euclidean distance between any two local point clouds to obtain a distance matrix, set the number of filtering scales and extract effective distances, obtain a set of thresholds using the quantile method, perform edge filtering based on radius thresholds, construct a higher-order graph, extract complete subgraphs to obtain topological objects, use the standard PH algorithm to obtain a set of barcodes and calculate the lifetime length to obtain a set of window topological features, and convert the barcode distance into a stability score. Based on the embedded coordinates and stability scores of discrete state sequences, feature vectors are constructed for each time point to generate a vector set. A scale set is set using linear sampling and a clustering algorithm is used to generate domain labels. The core score is calculated, and difference sequences are generated through homology dimension and stacked into a sequence matrix. The first principal component is extracted through PCA to obtain the topological kernel vector and construct the topological kernel matrix. The residual perturbation vector is obtained through singular value decomposition and the spectrum value is obtained through discrete Fourier transform. The split point index is calculated and combined with the stability score to obtain the depression risk.

2. The big data-based depression risk assessment method as described in claim 1, characterized in that: The difference between the computation node and its neighboring nodes is used to generate a behavior vector field. Embedded coordinates are generated by normalizing the Laplacian matrix through a random walk, including: A data matrix is ​​constructed for the discrete state sequence. The Euclidean distance between any two rows of the data matrix is ​​calculated using the Euclidean distance formula. The Euclidean distances are sorted in ascending order. The time points of the two rows corresponding to the smallest Euclidean distance are selected and marked as adjacency relationships. An adjacency graph is constructed. All adjacent points of each node are extracted. The difference between each node and all its corresponding adjacent points is calculated to obtain the state difference. The state difference is normalized to generate a behavior evolution direction vector. The behavior evolution direction vector is sorted according to the time points to obtain the behavior vector field. The edge weights are defined as rows, and the state vectors are defined as columns. A similarity matrix is ​​constructed, and a graph Laplacian matrix is ​​constructed using the random walk normalized Laplacian matrix method. The graph Laplacian matrix is ​​then decomposed into eigenvalues ​​to obtain the feature components and generate the embedded coordinates.

3. The big data-based depression risk assessment method as described in claim 2, characterized in that: The process of obtaining a threshold set through the quantile method, performing edge filtering based on the radius threshold, and constructing a higher-order graph includes: Construct a set of time windows, extract the embedding coordinates of the popular point clouds of the analysis time windows in the set of time windows, and obtain a set of local point clouds; For any two local point clouds in the local point cloud set, calculate the Euclidean distance using the Euclidean distance formula and generate a distance matrix; The number of filtering scales is set based on empirical rules. Effective distances are extracted based on the distance matrix. The starting and ending radius values ​​are calculated using the quantile method. Linear interpolation is performed on the starting and ending radius values ​​to obtain a threshold set. For each radius threshold in the threshold set, the local point cloud in the distance matrix is ​​traversed to construct a high-order graph.

4. The big data-based depression risk assessment method as described in claim 3, characterized in that: The process of converting barcode distance into a stability score includes: Extract complete subgraphs from higher-order graphs, generate topological objects, combine the Euclidean norm values ​​of the behavior vector field, and use the standard PH algorithm to calculate the barcode set in each homology dimension, including 0-dimensional, 1-dimensional, and 2-dimensional, and calculate the barcode weights. The barcode weights of each homology dimension in each analysis time window are summed to obtain the stability contribution value of each window. The stability contribution values ​​of all homology dimensions are weighted and summed to obtain the topological stability index and the window topological feature set. The distance between the window topological feature set and the healthy topological feature set is calculated using the barcode distance formula and converted into a stability score using the natural exponential function.

5. The big data-based depression risk assessment method as described in claim 4, characterized in that: The process of using linear sampling to define the scale set and clustering algorithms to generate domain labels, and calculating the core score, includes: A feature vector is constructed for each time point to obtain a vector set. A linear sampling method is used to set the scale parameter. For each scale parameter, a clustering algorithm is used to cluster the vector set to obtain the domain label. For each time point, iterate through each scale parameter in the scale set, select the domain label that appears most frequently, set it as the structural domain label, and calculate the core score.

6. The method for assessing depression risk based on big data as described in claim 5, characterized in that: The calculation of the split point index, combined with the stability score, yields the risk of depression, including: Based on the number of filtering scales, the homology dimensions under each filtering scale are arranged vertically to obtain curve vectors. First-order difference is performed on the curve vectors to obtain sequence matrices. Principal component analysis is performed on the matrix to extract the first principal component, which is set as the topological kernel vector. A topological kernel matrix is ​​constructed, and singular value decomposition is performed on the topological kernel matrix to calculate the residual perturbation vector. The discrete Fourier transform of the residual perturbation vector is used to obtain the spectral value, the split point index is calculated, and the overall structural stability is calculated in combination with the stability score. The overall structural stability is normalized to obtain the stability trajectory, and the risk of depression is calculated.

7. A big data-based depression risk assessment system, based on the big data-based depression risk assessment method according to any one of claims 1 to 6, characterized in that: include, The multimodal acquisition and preprocessing module is used to collect multimodal data through the API interface and generate preprocessed multimodal data using sliding window analysis and forward filling. The discrete state sequence construction module is used to construct time-indexed sequences by horizontally stacking and normalizing preprocessed multimodal data to generate discrete state sequences. The graph embedding and behavior vector field generation module is used to construct a data matrix based on discrete state sequences, generate an adjacency graph, a similarity matrix, embedding coordinates, and a behavior vector field, and obtain a popular point cloud. The Time Window Topology Analysis and Stability Scoring module is used to construct a set of time windows and generate topology objects, and to calculate the barcode set, barcode weight, topology stability index and stability score; The multi-scale structural domain identification and depression risk calculation module is used to construct feature vectors and generate domain labels, and calculate core degree score, topological kernel vector, split point index and comprehensive structural stability to obtain depression risk.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the big data-based depression risk assessment method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the big data-based depression risk assessment method according to any one of claims 1 to 6.