Biphase affective disorder early recognition method based on big data model

By combining individual dynamic gravitational field models and group behavior mapping models, the problem of the inability of existing technologies to capture the cohesive and progressive 'drift' of individual behavior patterns has been solved, enabling efficient early identification of bipolar disorder and improving the accuracy and reliability of identification.

CN122050841APending Publication Date: 2026-05-15XIAMEN XIANYUE HOSPITAL (XIAMEN MENTAL HEALTH CENT XIANYUE HOSPITAL AFFILIATED TO XIAMEN MEDICAL COLLEGE)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIAMEN XIANYUE HOSPITAL (XIAMEN MENTAL HEALTH CENT XIANYUE HOSPITAL AFFILIATED TO XIAMEN MEDICAL COLLEGE)
Filing Date
2026-04-02
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies cannot effectively capture the dynamic process of cohesive and gradual 'drift' in individual behavioral patterns, and lack collaborative quantitative analysis of the dynamic relationship between individual behavioral evolution and group reference system. This leads to the early identification of bipolar disorder being easily confused with pathological remodeling and normal fluctuations.

Method used

By constructing an individual dynamic gravitational field model and a group behavior map model, the gradual evolution of the cluster center of individual behavior patterns and the deviation of group behavior patterns are tracked respectively, generating a sequence of behavior gravitational intensity values ​​and a sequence of group pattern deviation values. These sequences are analyzed collaboratively to locate the 'system drift' characteristics, and a comprehensive judgment is made in combination with historical behavior data.

Benefits of technology

It significantly improves the accuracy and reliability of early identification of bipolar disorder, effectively distinguishes between pathological behavioral system evolution and ordinary behavioral fluctuations, and reduces misjudgment and missed diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122050841A_ABST
    Figure CN122050841A_ABST
Patent Text Reader

Abstract

The invention discloses a biphasic affective disorder early recognition method based on a big data model, and belongs to the technical field of mental health, and the method specifically comprises the steps: collecting multi-source behavior data of a target individual, and generating a time-synchronized behavior data unit through cross-modal alignment; inputting the behavior data unit into an individual dynamic gravitational field model and a group behavior map model in parallel, and respectively outputting a behavior gravitational intensity value sequence and a group pattern deviation value sequence; positioning an overlapping time interval in which the behavior gravitation intensity value is continuously increased and the group pattern deviation value is continuously decreased or unchanged in the two sequences; a rising slope sequence and a changing slope sequence are calculated based on data in the interval, and a fusion drift index is subjected to weighted synthesis; and finally, performing comprehensive judgment in combination with hierarchical features of the historical behavior pattern, and outputting an early risk identification result. According to the invention, automatic identification of the early risk of the bipolar affective disorder is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mental health technology, specifically to a method for early identification of bipolar disorder based on big data models. Background Technology

[0002] With the rapid development of digital health monitoring technology, the quantitative assessment of mental and psychological states using multi-source behavioral data generated by mobile terminals and wearable devices has become an important research direction. Bipolar disorder, as a complex and fluctuating mental illness, requires early identification for intervention, treatment, and prognosis improvement. Big data-based behavioral analysis offers a potential technological approach for capturing subtle and persistent behavioral pattern changes during its prodromal phase.

[0003] In existing technologies, researchers have attempted to achieve early risk warnings by collecting individuals' behavioral and physiological time-series data and using statistical models or machine learning methods for anomaly detection. Some approaches focus on comparing an individual's real-time behavioral data with a pre-established general healthy group behavioral benchmark, identifying anomalies by detecting significant deviations; others aim to construct an individual's historical behavioral baseline and detect risks by monitoring abrupt changes in current behavior relative to this baseline pattern. These methods have shown some effectiveness in identifying obvious and drastic behavioral anomalies. However, existing methods for identifying mental states based on behavioral data often focus on monitoring the absolute values ​​of individual behavioral characteristics or the degree of deviation from fixed patterns. They typically compare real-time individual data with a static, general healthy group benchmark or the individual's own fixed historical baseline. However, the behavioral pattern changes characteristic of early-stage bipolar disorder often manifest as a slow, cohesive "system drift," where the individual gradually forms a new, self-consistent but unhealthy behavioral pattern system, rather than simply violating their original habits or group norms. Traditional methods struggle to effectively distinguish this internally consistent behavioral pattern reconstruction from ordinary daily behavioral fluctuations or transient emotional reactions, easily leading to misjudgments or omissions. The fundamental reason is that current technologies lack a synergistic analysis of the internal evolutionary dynamics of the individual's behavioral system and its dynamic relationship with the group reference frame. They have not yet effectively quantified this gradual "drift" process or established a joint discriminative mechanism with changes in the group's behavioral space. Summary of the Invention

[0004] The purpose of this invention is to provide an early identification method for bipolar disorder based on a big data model, and to solve the following technical problems: Existing technologies rely on deviation detection from static benchmarks, which cannot effectively capture the dynamic process of cohesive and gradual "drift" in individual behavioral patterns. They also lack collaborative quantitative analysis of the dynamic relationship between individual behavioral evolution and group reference system, and are prone to confusing pathological remodeling with normal fluctuations.

[0005] The objective of this invention can be achieved through the following technical solutions: An early identification method for bipolar disorder based on big data models includes the following steps: S1. Continuously collect mobile terminal operation behavior data, wearable device physiological data and environmental audio data of target individuals to form multi-source heterogeneous raw behavioral time-series data; S2. Perform cross-modal slicing and time alignment on the original behavioral time series data to generate time-synchronized behavioral data units. Each behavioral data unit contains synchronized observations from all data sources within a fixed duration. S3. Input the behavioral data units into the individual dynamic gravitational field model in chronological order, calculate the distance between the newly arrived behavioral data units and the dynamically updated cluster center of individual behavioral patterns, and output the sequence of behavioral gravitational intensity values. S4. Input the behavioral data unit into the group behavior map model, calculate the deviation between the projection position of the behavioral data unit in the pre-constructed group behavior pattern space and the reference prototype, and output the group pattern deviation value sequence. S5. Traverse the sequence of behavioral gravity intensity values ​​and the sequence of group pattern deviation values ​​on the time axis to locate the overlapping time intervals where the behavioral gravity intensity values ​​continuously increase while the group pattern deviation values ​​continuously decrease or remain unchanged. S6. Extract the numerical calculation rising slope sequence of the behavior gravity intensity value sequence within the overlapping time interval, and the numerical calculation changing slope sequence of the group pattern deviation value sequence. Perform a weighted synthesis operation on the rising slope sequence and the changing slope sequence to generate the fusion drift index. S7. The drift index and the hierarchical features of behavioral patterns generated based on historical behavioral data units are combined to make a comprehensive judgment and output the early risk identification results of bipolar disorder.

[0006] As a further aspect of the present invention: the specific process of generating time-synchronized behavioral data units in S2 is as follows: The mobile terminal operation behavior data, wearable device physiological data and environmental audio data are timestamped and normalized respectively. The time base of the mobile terminal operation behavior data, wearable device physiological data and environmental audio data is unified. On the unified time axis, continuous and non-overlapping time windows are divided in units of fixed duration. For each time window, application switching frequency, average screen brightness, and number of touch events are extracted from mobile terminal operation behavior data; average heart rate and median skin conductivity are extracted from wearable device physiological data; and average decibel value and low-frequency energy ratio are extracted from environmental audio data. All observations extracted from mobile terminal operation behavior data, wearable device physiological data, and environmental audio data within the same time window are arranged in a preset order and combined to form the behavior data unit corresponding to the time window.

[0007] As a further aspect of the present invention: the specific process of outputting the sequence of gravitational intensity values ​​in step S3 is as follows: The individual dynamic gravitational field model stores cluster center vectors and update coefficients for individual behavior patterns. The observed value vector of the first behavior data unit is compared with the cluster center vector of the individual behavior pattern stored in the individual dynamic gravitational field model using Euclidean distance calculation. The reciprocal of the calculated distance value is recorded as the gravitational intensity value of the first behavior. Based on the update coefficients stored in the individual dynamic gravitational field model, the observed value vector of the first behavioral data unit and the cluster center vector of the individual behavioral pattern are weighted and averaged to obtain the updated cluster center vector of the individual behavioral pattern. For each subsequent behavioral data unit input in chronological order, the distance calculation, behavioral gravitational intensity value recording and cluster center vector update operations are performed sequentially to form a sequence of behavioral gravitational intensity values.

[0008] As a further aspect of the present invention: the construction process of the individual dynamic gravitational field model is as follows: All historical behavioral data units of the target individual within the historical baseline period are obtained to form an individual baseline dataset. The arithmetic mean of the observations of each dimension of all behavioral data units in the individual baseline dataset is calculated to generate a mean vector. The mean vector is set as the initial value of the cluster center vector of the individual behavioral pattern. Calculate the variance of the simulated sequence of behavioral gravity intensity corresponding to the individual baseline dataset. Based on the range of the variance value, select the corresponding value from the preset mapping table as the update coefficient. Associate and store the initialized individual behavioral pattern cluster center vector with the selected update coefficient to complete the construction of the individual dynamic gravity field model.

[0009] As a further aspect of the present invention: the specific process of outputting the group pattern deviation value sequence in S4 is as follows: The group behavior mapping model includes a set of projection rules for the group behavior pattern space and a set of coordinates for the group behavior reference prototypes. The observation vectors of the behavior data units are mapped to the group behavior pattern space according to the projection rules included in the group behavior mapping model to obtain the corresponding projection coordinates. The Euclidean distance between the projection coordinates and the coordinates of each group behavior reference prototype in the group behavior mapping model is calculated. The minimum value among all distances is selected as the original deviation value of the behavior data unit. The original deviation values ​​generated in time sequence are averaged using a sliding window, with the window length associated with a fixed duration of the behavior data unit. The processed group pattern deviation value sequence is then output.

[0010] As a further aspect of the present invention: the construction process of the group behavior graph model is as follows: A set of historical behavioral data units of healthy group samples is obtained to form a group training dataset. Principal component analysis is performed on the group training dataset to retain the first few principal component directions whose cumulative variance contribution rate exceeds a preset threshold. These principal component directions are used as basis vectors to span a reduced-dimensional subspace to form a projection rule. All samples in the group training dataset are projected onto the reduced-dimensional subspace. Cluster analysis is performed on the projected sample points. The coordinates of the center points of each cluster are determined as the coordinates of the group behavior reference prototype. The projection rules are associated with and stored with the coordinates of the group behavior reference prototype to complete the construction of the group behavior graph model.

[0011] As a further aspect of the present invention: in step S5, the process of locating the overlapping time interval is as follows: Set a minimum continuous point value, and start from the first time point to sequentially scan the sequence of gravitational intensity values. Identify and record all segments where the values ​​are continuously and strictly monotonically increasing and the number of increasing points is not less than the minimum continuous point value, and obtain the first segment set. The synchronous scanning group pattern deviation value sequence identifies and records all segments whose values ​​are continuously and strictly monotonically decreasing or whose values ​​are constant and whose number of consecutive points is not less than the minimum consecutive point value, thus obtaining a second segment set. Each segment in the first segment set is compared with all segments in the second segment set on the time axis to find all segment pairs with overlapping time parts. Each overlapping part is defined as an overlapping time interval.

[0012] As a further aspect of the present invention: in step S6, the process of generating the fusion drift index is as follows: For each overlapping time interval, the linear regression method is used to fit the sequence data points of the behavior gravity intensity value and the sequence data points of the group pattern deviation value within the overlapping time interval, respectively, to obtain the slope of the corresponding regression line, which is denoted as the interval rising slope and the interval change slope, respectively. The interval rising slopes of all overlapping time intervals are collected to form the rising slope sequence. The slope changes of all overlapping time intervals are collected to form a slope change sequence. A first weighting factor is assigned to each element in the rising slope sequence, and a second weighting factor is assigned to each element in the changing slope sequence. The weighted rising slope and the weighted changing slope of the same overlapping time interval are added together to obtain the fusion contribution value of the overlapping time interval. The arithmetic mean of all fusion contribution values ​​is calculated, and the arithmetic mean is linearly normalized to output the fusion drift index.

[0013] As a further aspect of the present invention: the specific process of outputting the early risk identification result of bipolar disorder in S7 is as follows: Predefine the positive judgment threshold of the fusion drift index, predefine the positive judgment threshold of the fluctuation feature in the hierarchical feature of the behavior pattern, read the currently calculated fusion drift index, and determine whether the fusion drift index reaches or exceeds the positive judgment threshold of the fusion drift index. Simultaneously, the system extracts the most recent consecutive time-layered fluctuation feature subsequences from the current behavioral pattern layered features, calculates the average value of the fluctuation feature subsequences, and determines whether the average value reaches or exceeds the positive judgment threshold of the fluctuation feature. When both the fusion drift index and the average value of the fluctuation feature reach or exceed their respective positive judgment thresholds, the system outputs a judgment result indicating the existence of early risk; when at least one of the fusion drift index and the average value of the fluctuation feature fails to reach its corresponding positive judgment threshold, the system outputs a judgment result indicating the absence of early risk.

[0014] The beneficial effects of this invention are: This invention dynamically tracks the gradual evolution of individual behavioral pattern cluster centers by constructing an individual dynamic gravitational field model, generating a sequence of behavioral gravitational strength values ​​characterizing the cohesion of behavior. Simultaneously, it uses a group behavior mapping model to calculate the relative deviation of behavioral data in the group behavior space, forming a group pattern deviation value sequence. Through collaborative analysis of these two sequences, it identifies overlapping time intervals where behavioral gravitational strength continuously increases while group deviation synchronously decreases or stabilizes. Based on this, a fusion drift index capable of quantifying "system drift" characteristics is synthesized. This method achieves, for the first time, dynamic monitoring of the cohesive reconstruction process of individual behavioral patterns and effectively distinguishes between pathological behavioral system evolution and ordinary behavioral fluctuations through dynamic comparison with a group reference system. Finally, it combines behavioral pattern hierarchical features extracted from historical behavioral data for comprehensive judgment, thereby significantly improving the accuracy and reliability of early identification of bipolar disorder and solving the problems of misjudgment and missed judgment caused by the reliance on static comparison in existing technologies. Attached Figure Description

[0015] The invention will now be further described with reference to the accompanying drawings.

[0016] Figure 1 This is a flowchart illustrating the present invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Please see Figure 1 As shown, this invention is an early identification method for bipolar disorder based on a big data model, comprising the following steps: S1. Continuously collect mobile terminal operation behavior data, wearable device physiological data and environmental audio data of target individuals to form multi-source heterogeneous raw behavioral time-series data; S2. Perform cross-modal slicing and time alignment on the original behavioral time series data to generate time-synchronized behavioral data units. Each behavioral data unit contains synchronized observations from all data sources within a fixed duration. S3. Input the behavioral data units into the individual dynamic gravitational field model in chronological order, calculate the distance between the newly arrived behavioral data units and the dynamically updated cluster center of individual behavioral patterns, and output the sequence of behavioral gravitational intensity values. S4. Input the behavioral data unit into the group behavior map model, calculate the deviation between the projection position of the behavioral data unit in the pre-constructed group behavior pattern space and the reference prototype, and output the group pattern deviation value sequence. S5. Traverse the sequence of behavioral gravity intensity values ​​and the sequence of group pattern deviation values ​​on the time axis to locate the overlapping time intervals where the behavioral gravity intensity values ​​continuously increase while the group pattern deviation values ​​continuously decrease or remain unchanged. S6. Extract the numerical calculation rising slope sequence of the behavior gravity intensity value sequence within the overlapping time interval, and the numerical calculation changing slope sequence of the group pattern deviation value sequence. Perform a weighted synthesis operation on the rising slope sequence and the changing slope sequence to generate the fusion drift index. S7. The drift index and the hierarchical features of behavioral patterns generated based on historical behavioral data units are combined to make a comprehensive judgment and output the early risk identification results of bipolar disorder.

[0019] In a preferred embodiment of the present invention, the specific process of generating time-synchronized behavioral data units in step S2 is as follows: First, the raw streaming data from three independent data sources undergoes timestamp normalization. Mobile terminal operational behavior data originates from the operating system event logs of smartphones or tablets, and its timestamps are typically based on the device's local system time. Wearable device physiological data comes from smartwatches or specialized biosensors, and its timestamps may be based on the device's internal clock or Coordinated Universal Time (UTC). Ambient audio data originates from periodic sampling records from the terminal's microphone, and its timestamps are typically synchronized with the audio hardware-driven clock. The core of the normalization process is to uniformly convert all timestamps to an absolute timeline based on UTC and compensate for potential network transmission latency and device clock drift. For example, by deploying a time synchronization daemon on each data acquisition terminal, this process periodically communicates with a network time protocol server and calibrates the local clock, ensuring that the deviation of all generated event timestamps is within 50 milliseconds.

[0020] After time synchronization is completed, continuous and non-overlapping time windows are divided on a unified absolute timeline, using fixed durations as the basic unit of slicing. This fixed duration can be set according to monitoring needs, for example, 30 minutes. This means that the first time window is defined as starting at 0 minutes 0 seconds and ending at 30 minutes 0 seconds, the second time window is defined as ending at 30 minutes 0 seconds and ending at 60 minutes 0 seconds, and so on.

[0021] For each time window, representative observations were extracted from three types of data. Observations extracted from mobile terminal operation behavior data included: application switching frequency, defined as the total number of times a user switched between different applications within the time window (one switch refers to jumping from foreground application A to foreground application B); average screen brightness, obtained by collecting screen brightness sensor values ​​once per second and calculating the arithmetic mean of all sampled values ​​within the window; and the number of touch events, referring to the total number of screen touch or swipe events recorded by the system within the time window.

[0022] The observations extracted from the physiological data of wearable devices include: average heart rate, which is calculated by arithmetically averaging the pulse values ​​reported per second by the photoelectric heart rate sensor within the time window; and median skin conductivity, which is the value in the middle position after sorting all the original skin conductivity samples collected by the skin conductance sensor within the time window by size. The median is used to reduce the interference of extreme values ​​caused by factors such as instantaneous sweat secretion.

[0023] The observations extracted from environmental audio data include: the average decibel value, which is obtained by arithmetically averaging the sound pressure level decibel values ​​calculated for every 100 milliseconds of audio frames within the time window; and the low-frequency energy ratio, which is calculated by first performing a fast Fourier transform on all audio frames within the window to obtain the spectrum, then calculating the sum of the energy values ​​of all frequency points in the frequency range of 85 Hz to 255 Hz, and finally dividing the sum by the total energy of the entire frequency band from 20 Hz to 8000 Hz. The resulting ratio is the low-frequency energy ratio, which can be used to indirectly reflect whether there is continuous low-frequency speech or noise in the environment.

[0024] After extracting the above observations, they are arranged and combined in a preset, fixed order. For example, a possible order is: application switching frequency, average screen brightness, number of touch events, average heart rate, median skin conductivity, average decibel value, and low-frequency energy percentage. These seven observations arranged in this order are combined to form a one-dimensional vector, which represents a complete behavioral data unit corresponding to that specific time window. The behavioral data unit is the core data structure defined in this scheme. It transforms the multi-source, heterogeneous raw data stream into a unified, fixed-dimensional, time-aligned numerical representation, providing standardized input for subsequent model calculations.

[0025] Acquiring mobile terminal operation behavior data, wearable device physiological data, and environmental audio data is based on multimodal fusion analysis. Mobile terminal operation behavior data directly reflects the user's proactive behavior patterns, attention allocation, and sleep patterns. Wearable device physiological data provides objective physiological indicators of the user's autonomic nervous system activity and emotional arousal. Environmental audio data can capture the user's social interactions, the stress level of the environment, or whether they are in a state of social isolation. These three types of data characterize an individual's state from three different and complementary dimensions: behavior, physiology, and environment. Their simultaneous collection and fusion analysis can construct a more comprehensive and robust psychological state assessment model than a single data type, which is particularly important for identifying the complex and multidimensional characteristics that may accompany the early stages of bipolar disorder, such as social withdrawal, changes in activity levels, and abnormal emotional and physiological responses.

[0026] In another preferred embodiment of the present invention, the specific process of outputting the sequence of gravitational intensity values ​​in step S3 is as follows: In this embodiment, a behavioral data unit is a vector containing seven observations, as described above. The observation vector refers to the numerical representation of this behavioral data unit. The individual behavioral pattern cluster center vector is a vector with the same dimension as the observation vector, also a seven-dimensional vector in this example. It represents a mathematical abstraction of the typical or average behavioral pattern of the individual over a certain period of time. The update coefficient is a decimal between 0 and 1, used to control the speed at which the individual behavioral pattern cluster center vector tracks the latest behavioral data unit observation vector. The behavioral gravity strength value is a scalar value, whose physical meaning can be understood as the similarity or attractive force between the current behavioral data unit and the current individual behavioral pattern cluster center vector; a higher value indicates greater similarity.

[0027] The construction of the individual dynamic gravitational field model is completed offline. First, it is necessary to acquire all historical behavioral data units of the target individual within a historical baseline period. The historical baseline period should be selected when the individual is in a stable psychological state with no significant abnormal reports, for example, a continuous 30 days. All behavioral data units generated within these 30 days at fixed time windows constitute the individual baseline dataset. Assuming a time window of 30 minutes, 48 ​​behavioral data units are generated per day, resulting in 1440 behavioral data units over 30 days.

[0028] Next, the arithmetic mean of the observations for each dimension of all behavioral data units in the individual baseline dataset is calculated to generate an average vector. Specifically, for the first dimension, application switching frequency, the application switching frequency values ​​of the 1440 behavioral data units are summed and then divided by 1440 to obtain the average application switching frequency. Similarly, the average screen brightness, average number of touch events, average heart rate, median average skin conductivity, average decibel level, and average low-frequency energy percentage are calculated. These seven averages are arranged in a fixed order, and the resulting 7-dimensional average vector is set as the initial value for the individual's behavioral pattern cluster center vector. This initial value represents the individual's baseline behavioral pattern during the stable period.

[0029] Next, the update coefficients need to be determined. This step calculates the variance of a simulated sequence of behavioral gravity intensity by simulating the operation of an individual dynamic gravitational field model, and determines the update coefficients based on this variance. The simulation process is as follows: Starting with the initial value of the individual behavioral pattern cluster center vector just calculated, historical behavioral data units from the individual baseline dataset are input into a simulation calculation process one by one according to their original chronological order. For the first input historical behavioral data unit, the Euclidean distance between its observed value vector and the initial value of the current individual behavioral pattern cluster center vector is calculated.

[0030] The Euclidean distance is calculated as the square root of the sum of the squares of the differences in the corresponding dimensions of the two vectors. The calculated distance value is denoted as d1. Next, the first simulated value of the behavioral gravity intensity is calculated as 1 divided by d1 plus a very small constant epsilon, for example, 0.0001. Epsilon is added to prevent division by zero when the distance is zero. This simulated value is recorded. Then, the cluster center vector of the individual behavioral pattern is updated based on a provisional update coefficient, for example, 0.01. The update formula is: the new cluster center vector equals the current cluster center vector multiplied by 1 minus the update coefficient, plus the observed value vector of the currently input behavioral data unit multiplied by the update coefficient. This operation is a weighted average calculation. After the update, the above distance calculation, behavioral gravity intensity simulation value recording, and cluster center vector update operations are repeated with the next historical behavioral data unit until all 1440 historical behavioral data units in the individual baseline dataset have been processed, ultimately resulting in a behavioral gravity intensity simulation sequence containing 1440 values.

[0031] Calculate the variance of the simulated sequence of behavioral gravity intensity. Variance is a statistic that measures the degree of fluctuation in the sequence. Since the choice of update coefficients directly affects the update speed of the cluster center vector, thus affecting the smoothness and volatility of the behavioral gravity intensity value sequence, it is necessary to determine the appropriate update coefficients in reverse based on the desired level of sequence volatility. For this purpose, we pre-define a mapping table. For example, when the simulated sequence variance is less than 0.05, it indicates that the individual baseline behavior is very stable, and the update coefficient can be set larger to accelerate the tracking of potential changes; the mapping value is 0.02. When the simulated sequence variance is between 0.05 and 0.2, the update coefficient is mapped to 0.01. When the simulated sequence variance is greater than 0.2, it indicates that the baseline behavior itself has some volatility; the update coefficient should be set smaller to maintain cluster center stability; the mapping value is 0.005. By consulting this pre-defined mapping table, the final update coefficients can be determined based on the calculated simulated sequence variance.

[0032] Finally, the initialized individual behavior pattern cluster center vectors are associated with and stored with the determined update coefficients, for example, in a specific model configuration file. This completes the construction of the individual dynamic gravitational field model.

[0033] During the model deployment phase, the specific process for outputting the behavioral gravity intensity value sequence is as follows: First, the pre-constructed individual dynamic gravity field model is loaded, and the initial values ​​and update coefficients of the cluster center vectors of individual behavioral patterns stored within it are read. When real-time monitoring begins, the first time window ends, generating the first real-time behavioral data unit. The Euclidean distance between its observed value vector and the loaded individual behavioral pattern cluster center vector is calculated, yielding the distance value dreal1. Subsequently, the first true behavioral gravity intensity value is calculated, which is 1 divided by dreal1 plus the minimal constant epsilon. This value is recorded as the first point in the behavioral gravity intensity value sequence.

[0034] Next, based on the update coefficient stored in the model, for example, 0.01, the cluster center vector of the individual behavior pattern is updated for the first time. The update operation is the same as the simulation process, i.e., a weighted average is performed. Assuming the update coefficient is 0.01, the new cluster center vector is equal to the current cluster center vector multiplied by 0.99, plus the observation vector of the first real-time behavior data unit multiplied by 0.01. This operation fine-tunes the cluster center vector to reflect the most recently observed behavior pattern.

[0035] When the second time window ends and the second real-time behavior data unit is generated, the above process is repeated: the Euclidean distance is calculated between the updated individual behavior pattern cluster center vector and the second observation value vector, and then the second behavior gravity intensity value is calculated and recorded. Then, the cluster center vector is updated with a new round of weighted average using the second observation value vector and the update coefficient.

[0036] This process continues, and for each newly arrived real-time behavioral data unit in chronological order, the following three operations are performed sequentially: calculate the Euclidean distance between its observation vector and the current individual behavioral pattern cluster center vector; convert the distance value to its reciprocal and record it as the behavioral gravity intensity value at the current moment; and update the individual behavioral pattern cluster center vector by weighted averaging using the current observation vector and a fixed update coefficient.

[0037] Through this iterative mechanism, the cluster center vector of an individual behavioral pattern becomes a dynamically moving reference point, slowly tracking the long-term changing trend of the individual's behavioral pattern. The sequence of behavioral gravity intensity values ​​reflects the instantaneous similarity between each new behavioral data unit and this dynamically changing individual normal pattern. If an individual's behavioral pattern undergoes a slow, consistent change, the dynamically updated cluster center will gradually follow this drift, and the distance to the new behavioral data unit, i.e., the behavioral gravity intensity value, will exhibit a specific pattern of change. For example, when a behavioral pattern undergoes cohesive reconstruction, although the new behavioral pattern deviates from the old baseline, because the cluster center is constantly being updated and adjusted, the distance between the new behavioral data unit and the dynamic cluster center may remain small or even decrease, resulting in an upward trend in the behavioral gravity intensity value. This is precisely the core dynamic feature that this embodiment aims to capture. The entire process is based entirely on mathematical calculations, requiring no subjective judgment, and achieves automated and quantitative tracking of the intrinsic evolutionary dynamics of individual behavioral patterns.

[0038] In another preferred embodiment of the present invention, the specific process of outputting the group pattern deviation value sequence in step S4 is as follows: The group behavior pattern space is a low-dimensional real-number vector space constructed mathematically, with a dimension much lower than that of the original behavioral data units. The design goal of this space is to preserve the most significant differences in the behavioral patterns of healthy groups. The projection rule is a set of mathematical transformation coefficients used to map the observation vectors of high-dimensional behavioral data units to the low-dimensional group behavior pattern space. The group behavior reference prototype is a series of representative coordinate points in the group behavior pattern space, each point corresponding to a typical healthy group behavior pattern. The original deviation value is a scalar representing the minimum geometric distance between the position of a specific behavioral data unit after projection and the typical pattern of the entire healthy group. The group pattern deviation value sequence is a time series obtained by temporally smoothing the original deviation values, used to reflect the degree of deviation of individual behavior from the normal state of the healthy group over time.

[0039] The construction of the group behavior mapping model is an offline training process. First, it's necessary to obtain a set of historical behavioral data units from a healthy group sample to form the group training dataset. The selection criteria for the healthy group include, but are not limited to, the absence of clinical diagnoses of mental illness and self-reported absence of recent major life stressors. For example, behavioral data units from 1000 healthy volunteers over the past 90 days can be obtained from collaborating research institutions. These data units are generated according to the same specifications as the target individuals, i.e., one unit every 30 minutes, 48 ​​units per day. After removing samples with severely missing data, assuming the remaining valid sample size is 950 people, each providing 90 days of data, the group training dataset contains a total of 950 x 90 x 48 = 4,104,000 behavioral data units. Each data unit is a vector containing 7 observations.

[0040] Next, principal component analysis (PCA) is performed on the training dataset to achieve dimensionality reduction and form projection rules. PCA is a statistical method used to identify the orthogonal axes with the largest directions of variation in the data. Specifically, the 4,104,000 behavioral data units are first treated as a 4,104,000-row, 7-column matrix, and the covariance matrix across its seven dimensions is calculated. Then, the eigenvalues ​​and eigenvectors of this covariance matrix are calculated. The magnitude of the eigenvalue represents the magnitude of the data variance along the corresponding eigenvector direction. The eigenvalues ​​are sorted from largest to smallest, and the cumulative variance contribution rate is calculated. The cumulative variance contribution rate is the percentage of the sum of the top N largest eigenvalues ​​relative to the sum of all eigenvalues. A threshold is preset, for example, 85%. This means that we want to retain principal component directions that can explain 85% of the variability in the original data. Suppose that after calculation, the cumulative contribution rate of the top 3 eigenvalues ​​reaches 87%, exceeding the 85% threshold, then we retain these top 3 principal component directions. These 3 eigenvectors are each a 7-dimensional vector, and they are orthogonal to each other. Using these three vectors as new coordinate basis vectors, the three-dimensional real space they span is the reduced-dimensional subspace we construct, which is the group behavior pattern space. The matrix formed by these three basis vectors is called the projection matrix. This projection matrix constitutes the core of the projection rule. Multiplying the observation vector of any 7-dimensional behavior data unit on the right by the transpose of this projection matrix yields its three-dimensional projected coordinates in the three-dimensional group behavior pattern space.

[0041] Subsequently, all 4,104,000 samples in the entire group training dataset were projected into this three-dimensional space, resulting in 4,104,000 three-dimensional coordinate points. Next, cluster analysis was performed on these projected points to identify naturally occurring behavioral pattern categories within the healthy group. The K-means clustering algorithm was used here. The number of clusters, K, is a parameter that needs to be determined. The clustering quality was evaluated by calculating the silhouette coefficient for different K values. The silhouette coefficient ranges from -1 to +1, with a larger value indicating better clustering. For example, calculations were performed with K equal to 3, 4, 5, and 6, and it was found that the silhouette coefficient was highest at 0.6 when K equaled 5. Therefore, the number of clusters was determined to be 5. After running the K-means algorithm, the 4,104,000 three-dimensional points were divided into 5 clusters, and the coordinates of the center point of each cluster were calculated. The three-dimensional coordinates of these 5 center points were then determined as the coordinates of 5 group behavior reference prototypes. Each prototype represents a stable combination of behavioral patterns that is prevalent in the healthy group.

[0042] Finally, the projection rules, i.e., the 3D projection matrix, are associated with and stored with the coordinates of the five group behavior reference prototypes, for example, by saving them as a data file. This completes the construction of the group behavior map model.

[0043] During the model deployment phase, the specific process for outputting the group pattern deviation value sequence is as follows: First, the pre-constructed group behavior map model is loaded, and the stored projection matrix and group behavior reference prototype coordinates are read. When a real-time behavior data unit of the target individual to be evaluated is input, such as a 7-dimensional observation vector named V, it is first mapped to the group behavior pattern space according to the projection rules. Specifically, the product of V and the transpose of the projection matrix is ​​calculated to obtain a 3-dimensional vector P, which is the projection coordinate of V in the group behavior pattern space.

[0044] Next, the Euclidean distance between the projected coordinates P and the coordinates of each group behavior reference prototype stored in the model is calculated. Assume the coordinates of the five prototypes are C1, C2, C3, C4, and C5. The distance D1 from P to C1, the distance D2 from P to C2, and so on up to D5 are calculated. The Euclidean distance is calculated as the square root of the sum of the squares of the differences in each dimension between two points in three-dimensional space. From these five distance values ​​D1 to D5, the smallest value is selected and denoted as Dmin. This Dmin is defined as the original deviation value of the behavioral data unit V. Its physical meaning is the difference between the current individual's behavioral pattern and the pattern closest to the typical pattern of all healthy groups. If Dmin is small, it indicates that the current behavior is very close to the healthy norm; if Dmin is large, it indicates that the current behavior deviates far from all healthy norms.

[0045] As the time window progresses, each new time window generates a new behavioral data unit. Repeating the projection and distance calculation steps above yields a sequence of raw deviation values ​​arranged chronologically. Assuming a time window of 30 minutes, then 48 raw deviation values ​​will be generated per day.

[0046] Using this raw deviation sequence directly may contain too much transient fluctuation or noise. Therefore, a sliding window averaging process is needed to smooth the sequence and highlight the trend. The window length needs to be related to the fixed duration of the behavioral data unit. A reasonable setting is for the window length to cover a one-day behavioral cycle, i.e., containing 48 data points. This means that when processing the t-th raw deviation, the first 48 raw deviations, including t, are taken; if there are fewer than 48, all existing values ​​are taken, and the arithmetic mean of these values ​​is calculated. This average is then used as the population pattern deviation output at time t.

[0047] For example, when processing the 49th raw deviation value R49, the average of the 48 values ​​from R2 to R49 is taken as the population pattern deviation value corresponding to the 49th time point. When processing the 50th raw deviation value R50, the average of the 48 values ​​from R3 to R50 is taken, and so on. This processing is equivalent to a moving average filter with a width of 48 units of time, which preserves the overall deviation trend over a daily period while smoothing out drastic fluctuations that may occur in a short period of time due to measurement errors or random behavior.

[0048] Ultimately, by continuously inputting real-time behavioral data units and sequentially performing projection, minimum distance calculation, and moving average, a smooth, time-synchronized sequence of group pattern deviation values ​​is formed. This sequence quantifies the dynamic deviation of the target individual's behavioral pattern relative to a healthy group reference frame over time. When co-analyzed with the behavioral gravity intensity value sequence generated by the individual dynamic gravity field model, it can provide an external reference perspective. For example, when the behavioral gravity intensity value increases, indicating that the individual's behavioral pattern is becoming more internally consistent and stable, if the group pattern deviation value decreases synchronously, it means that this newly formed stable pattern is approaching a healthy group prototype, which may be a positive adjustment. Conversely, if the behavioral gravity intensity value increases while the group pattern deviation value also increases or remains high, it suggests that the individual is forming a pattern that is both cohesive and stable but deviates from the healthy norm, which may be a risk signal requiring attention. This dual-sequence comparative analysis provides crucial evidence for early identification.

[0049] In another preferred embodiment of the present invention, the process of locating the overlapping time interval in step S5 is as follows: First, let's explain the process of locating overlapping time intervals in S5. An overlapping time interval is a key definition; it refers to a continuous period on the time axis where the behavioral gravity intensity value sequence exhibits a specific upward pattern, and the group pattern deviation value sequence exhibits a specific downward or stable pattern, with the two patterns completely or partially overlapping in time. The minimum consecutive point value is a preset integer parameter used to define the minimum number of consecutive data points necessary for a meaningful trend. Its purpose is to filter out short-lived and unreliable trend signals caused by noise or random fluctuations. For example, assuming a data sampling interval of 30 minutes, if the desired trend lasts at least 3 hours, the minimum consecutive point value should be set to 6, because 3 hours contains 6 30-minute time points.

[0050] The localization process begins by simultaneously traversing two generated time series: the behavioral gravity intensity value sequence and the group pattern deviation value sequence. These two sequences are strictly aligned in time. Starting from the first time point index 1, the behavioral gravity intensity value sequence is scanned sequentially. The goal of the scan is to identify all segments that satisfy the condition of strictly monotonically increasing values ​​with the number of consecutive increasing points not less than the minimum consecutive point value. Strict monotonically increasing means that the value of each subsequent data point in the sequence is strictly greater than the value of the previous data point. For example, for a numerical sequence segment [0.5, 0.52, 0.55, 0.57, 0.60, 0.63], if the minimum consecutive point value is 5, then the segment meets the condition. The identification algorithm records the start and end time point indices of each such segment, and all identified segments constitute the first segment set.

[0051] Simultaneously, the algorithm scans the sequence of deviation values ​​from the group pattern, identifying all segments that satisfy either a strictly monotonically decreasing numerical value or a constant numerical value with the number of consecutive points not less than the minimum consecutive point value. Strictly monotonically decreasing means that the value of each subsequent point is strictly less than the value of the previous point. A constant numerical value means that the values ​​of consecutive data points remain unchanged, with the variation range considered constant within a very small interval, such as 0.001. For example, one segment [1.2, 1.15, 1.1, 1.05, 1.0] satisfies the monotonically decreasing condition; another segment [0.8, 0.8, 0.8, 0.8, 0.8] satisfies the constant condition. Similarly, the algorithm records the start and end indices of each such segment, forming a second set of segments.

[0052] Then, a time comparison is performed. Take any segment F1 from the first segment set, with a time range from index StartA to EndA. Iterate through all segments in the second segment set, searching for a segment F2 that intersects with F1 on the time axis, with the time range of F2 being StartB to EndB. The condition for intersection is that the end index EndA of F1 is greater than or equal to the start index StartB of F2, and the start index StartA of F1 is less than or equal to the end index EndB of F2. If these conditions are met, the start index of the overlapping portion is the larger of StartA and StartB, and the end index is the smaller of EndA and EndB. This interval, defined by the larger start index and the smaller end index, is defined as an overlapping time interval. This comparison operation is repeated for each segment in the first segment set, ultimately identifying all overlapping time intervals that satisfy the collaborative change pattern. These intervals are the potential risk-related periods identified by this scheme.

[0053] In another preferred embodiment of the present invention, the process of generating the fusion drift index in step S6 is as follows: The first step is to calculate the slope of the trend of the two sequences within each identified overlapping time interval. Taking one overlapping time interval as an example, assume that the interval contains N data points from time index Tstart to Tend. For the behavioral gravity intensity value sequence, extract the values ​​of these N points and their corresponding time indices. Use linear regression to fit the data, aiming to find a straight line that minimizes the sum of the squared vertical distances from this line to all data points. The slope of this optimally fitted line, i.e., the rate of change of the behavioral gravity intensity value per unit time, is calculated and denoted as the interval rise slope, Kup. Similarly, perform linear regression on the N points of the population pattern deviation value sequence within the interval to obtain the slope of the optimally fitted line, denoted as the interval change slope Kchange. Since the behavioral gravity intensity value is expected to increase within the overlapping interval, Kup is usually positive; the population pattern deviation value is expected to decrease or stabilize, so Kchange is usually zero or negative.

[0054] The second step is to aggregate the calculation results for all overlapping time intervals. Assume a total of M overlapping time intervals are located. Arrange the Kup values ​​of all M intervals in order to form an ascending slope sequence Sup of length M. Similarly, arrange the Kchange values ​​of all M intervals in order to form a changing slope sequence Schange, also of length M.

[0055] The third step is weighted synthesis. A first weight factor W1 is assigned to each slope value in the rising slope sequence Sup, and a second weight factor W2 is assigned to each slope value in the changing slope sequence Schange. The weight factors can be set based on the length or significance of the corresponding overlapping time intervals. For example, a simple implementation is to make all weight factors equal to 1. More complex implementations can set weights based on the interval length N, with longer intervals assigned greater weights because trends over longer periods are considered more reliable. Assuming equal weights are used, then W1i = 1, W2i = 1, where i ranges from 1 to M.

[0056] The fourth step is to calculate the fusion contribution value for each interval. For the i-th overlapping time interval, the fusion contribution value Fi is calculated using the formula: Fi = W1i * Kupi + W2i * Kchangei. Since Kupi is positive and Kchangei is negative or zero, this operation essentially combines the positive gravitational change with the negative deviation change. If the two are significantly correlated, i.e., one is significantly positive and the other is significantly negative, then Fi will be a large positive value.

[0057] The fifth step is summarization and normalization. Calculate the arithmetic mean of all M fusion contribution values ​​Fi, denoted as AvgF. Then, perform linear normalization on AvgF, mapping it to a standard range, such as 0 to 100. Normalization requires determining the minimum and maximum values ​​based on historical data or theoretical boundaries. For example, by analyzing a large amount of historical health data, the fluctuation range of AvgF is determined to be roughly between L and U. The normalization index I is then calculated as: I = 100 * (AvgF - L) / (UL). If AvgF is less than L, I is 0; if it is greater than U, I is 100. This final I value is the fusion drift index. It is a number between 0 and 100; the higher the value, the stronger the observed cooperative behavior drift pattern.

[0058] In another preferred embodiment of the present invention, the specific process of outputting the early risk identification result of bipolar disorder in step S7 is as follows: This step requires two inputs: the most recently calculated fusion drift index I, and the hierarchical behavioral pattern features generated from the target individual's historical behavior. The hierarchical behavioral pattern features are a concept that needs to be explained in detail here, and in this embodiment, they are supplemented as follows: All historical behavioral data units of the target individual over a relatively long period, such as 90 days, are collected. These 90 days of data are divided into layers of 30 days each, resulting in three layers. For all behavioral data units within each layer, the standard deviation of its observation vector across each dimension is calculated, yielding a volatility feature vector characterizing the degree of behavioral volatility in that layer. Connecting the volatility feature vectors of these three layers in chronological order constitutes a hierarchical behavioral pattern feature describing the temporal variation of individual behavioral volatility. The volatility feature of the most recent layer reflects the recent behavioral stability.

[0059] The decision logic requires two preset positive decision thresholds. The first is the positive decision threshold ThetaI for the fused drift index, which can be set to 65, for example. The second is the positive decision threshold ThetaF for the volatility features in the hierarchical features of the behavioral pattern. For volatility features, we focus on the volatility of the most recent layer. Specifically, we extract the volatility feature subsequence of the most recent N consecutive time layers from the hierarchical features, for example, the most recent two layers. We calculate the average value of all volatility feature values ​​in this subsequence, denoted as AvgFluctuation. The threshold ThetaF may be a norm threshold of a multidimensional vector. In a simplified model, it can be a threshold of a key dimension of the volatility feature (such as the standard deviation of the applied switching frequency), for example, 0.15.

[0060] The decision process involves parallel checks. The first check reads the current fusion drift index I and checks if the condition I ≥ ThetaI is true. The second check reads the currently calculated AvgFluctuation and checks if the condition AvgFluctuation ≥ ThetaF is true. The final output is a binary decision. The decision logic outputs a signal representing early risk of bipolar disorder, such as a flag with a value of 1 or a "high risk" text label, only if both conditions are true. If either condition is false, it outputs a signal representing the absence of early risk, such as a flag with a value of 0 or a "low risk" text label. This AND logic ensures that a high-risk warning is triggered only when both a significant behavioral pattern co-drift signal and a recent increase in behavioral volatility are detected, improving the specificity of the decision and reducing false alarms. The entire process, from data preprocessing, dual-model calculation, co-drift localization, index synthesis to final decision, forms a complete and automated early risk identification technology solution.

[0061] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A method for early identification of bipolar disorder based on big data models, characterized in that, Includes the following steps: S1. Continuously collect mobile terminal operation behavior data, wearable device physiological data and environmental audio data of target individuals to form multi-source heterogeneous raw behavioral time-series data; S2. Perform cross-modal slicing and time alignment on the original behavioral time series data to generate time-synchronized behavioral data units. Each behavioral data unit contains synchronized observations from all data sources within a fixed duration. S3. Input the behavioral data units into the individual dynamic gravitational field model in chronological order, calculate the distance between the newly arrived behavioral data units and the dynamically updated cluster center of individual behavioral patterns, and output the sequence of behavioral gravitational intensity values. S4. Input the behavioral data unit into the group behavior map model, calculate the deviation between the projection position of the behavioral data unit in the pre-constructed group behavior pattern space and the reference prototype, and output the group pattern deviation value sequence. S5. Traverse the sequence of behavioral gravity intensity values ​​and the sequence of group pattern deviation values ​​on the time axis to locate the overlapping time intervals where the behavioral gravity intensity values ​​continuously increase while the group pattern deviation values ​​continuously decrease or remain unchanged. S6. Extract the numerical calculation rising slope sequence of the behavior gravity intensity value sequence within the overlapping time interval, and the numerical calculation changing slope sequence of the group pattern deviation value sequence. Perform a weighted synthesis operation on the rising slope sequence and the changing slope sequence to generate the fusion drift index. S7. The drift index and the hierarchical features of behavioral patterns generated based on historical behavioral data units are combined to make a comprehensive judgment and output the early risk identification results of bipolar disorder.

2. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In step S2, the specific process of generating time-synchronized behavioral data units is as follows: The mobile terminal operation behavior data, wearable device physiological data and environmental audio data are timestamped and normalized respectively. The time base of the mobile terminal operation behavior data, wearable device physiological data and environmental audio data is unified. On the unified time axis, continuous and non-overlapping time windows are divided in units of fixed duration. For each time window, application switching frequency, average screen brightness, and number of touch events are extracted from mobile terminal operation behavior data; average heart rate and median skin conductivity are extracted from wearable device physiological data; and average decibel value and low-frequency energy ratio are extracted from environmental audio data. All observations extracted from mobile terminal operation behavior data, wearable device physiological data, and environmental audio data within the same time window are arranged in a preset order and combined to form the behavior data unit corresponding to the time window.

3. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In S3, the specific process of outputting the sequence of gravitational intensity values ​​is as follows: The individual dynamic gravitational field model stores cluster center vectors and update coefficients for individual behavior patterns. The observed value vector of the first behavior data unit is compared with the cluster center vector of the individual behavior pattern stored in the individual dynamic gravitational field model using Euclidean distance calculation. The reciprocal of the calculated distance value is recorded as the gravitational intensity value of the first behavior. Based on the update coefficients stored in the individual dynamic gravitational field model, the observed value vector of the first behavioral data unit and the cluster center vector of the individual behavioral pattern are weighted and averaged to obtain the updated cluster center vector of the individual behavioral pattern. For each subsequent behavioral data unit input in chronological order, the distance calculation, behavioral gravitational intensity value recording and cluster center vector update operations are performed sequentially to form a sequence of behavioral gravitational intensity values.

4. The method for early identification of bipolar disorder based on a big data model according to claim 3, characterized in that, The construction process of the individual dynamic gravitational field model is as follows: All historical behavioral data units of the target individual within the historical baseline period are obtained to form an individual baseline dataset. The arithmetic mean of the observations of each dimension of all behavioral data units in the individual baseline dataset is calculated to generate a mean vector. The mean vector is set as the initial value of the cluster center vector of the individual behavioral pattern. Calculate the variance of the simulated sequence of behavioral gravity intensity corresponding to the individual baseline dataset. Based on the range of the variance value, select the corresponding value from the preset mapping table as the update coefficient. Associate and store the initialized individual behavioral pattern cluster center vector with the selected update coefficient to complete the construction of the individual dynamic gravity field model.

5. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In step S4, the specific process of outputting the population pattern deviation value sequence is as follows: The group behavior mapping model includes a set of projection rules for the group behavior pattern space and a set of coordinates for the group behavior reference prototypes. The observation vectors of the behavior data units are mapped to the group behavior pattern space according to the projection rules included in the group behavior mapping model to obtain the corresponding projection coordinates. The Euclidean distance between the projection coordinates and the coordinates of each group behavior reference prototype in the group behavior mapping model is calculated. The minimum value among all distances is selected as the original deviation value of the behavior data unit. The original deviation values ​​generated in time sequence are averaged using a sliding window, with the window length associated with a fixed duration of the behavior data unit. The processed group pattern deviation value sequence is then output.

6. The method for early identification of bipolar disorder based on a big data model according to claim 5, characterized in that, The construction process of the group behavior graph model is as follows: A set of historical behavioral data units of healthy group samples is obtained to form a group training dataset. Principal component analysis is performed on the group training dataset to retain the first few principal component directions whose cumulative variance contribution rate exceeds a preset threshold. These principal component directions are used as basis vectors to span a reduced-dimensional subspace to form a projection rule. All samples in the group training dataset are projected onto the reduced-dimensional subspace. Cluster analysis is performed on the projected sample points. The coordinates of the center points of each cluster are determined as the coordinates of the group behavior reference prototype. The projection rules are associated with and stored with the coordinates of the group behavior reference prototype to complete the construction of the group behavior graph model.

7. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In step S5, the process of locating the overlapping time interval is as follows: Set a minimum continuous point value, and start from the first time point to sequentially scan the sequence of gravitational intensity values. Identify and record all segments where the values ​​are continuously and strictly monotonically increasing and the number of increasing points is not less than the minimum continuous point value, and obtain the first segment set. The synchronous scanning group pattern deviation value sequence identifies and records all segments whose values ​​are continuously and strictly monotonically decreasing or whose values ​​are constant and whose number of consecutive points is not less than the minimum consecutive point value, thus obtaining a second segment set. Each segment in the first segment set is compared with all segments in the second segment set on the time axis to find all segment pairs with overlapping time parts. Each overlapping part is defined as an overlapping time interval.

8. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In step S6, the process of generating the fusion drift index is as follows: For each overlapping time interval, the linear regression method is used to fit the sequence data points of the behavior gravity intensity value and the sequence data points of the group pattern deviation value within the overlapping time interval, respectively, to obtain the slope of the corresponding regression line, which is denoted as the interval rising slope and the interval change slope, respectively. The interval rising slopes of all overlapping time intervals are collected to form the rising slope sequence. The slope changes of all overlapping time intervals are collected to form a slope change sequence. A first weighting factor is assigned to each element in the rising slope sequence, and a second weighting factor is assigned to each element in the changing slope sequence. The weighted rising slope and the weighted changing slope of the same overlapping time interval are added together to obtain the fusion contribution value of the overlapping time interval. The arithmetic mean of all fusion contribution values ​​is calculated, and the arithmetic mean is linearly normalized to output the fusion drift index.

9. The method for early identification of bipolar disorder based on a big data model according to claim 1, characterized in that, In step S7, the specific process for outputting the early risk identification results for bipolar disorder is as follows: Predefine the positive judgment threshold of the fusion drift index, predefine the positive judgment threshold of the fluctuation feature in the hierarchical feature of the behavior pattern, read the currently calculated fusion drift index, and determine whether the fusion drift index reaches or exceeds the positive judgment threshold of the fusion drift index. Simultaneously, extract the most recent consecutive time-layered fluctuation feature subsequences from the current behavior pattern layered features, calculate the average value of the fluctuation feature subsequences, and determine whether the average value reaches or exceeds the positive judgment threshold of the fluctuation feature. When both the fused drift index and the average value of the fluctuation feature reach or exceed their respective positive judgment thresholds, output a judgment result representing the existence of early risk. When at least one of the fusion drift index and the average volatility characteristic fails to reach the corresponding positive judgment threshold, the output indicates that there is no early risk.