Multimodal Feature Fusion and Recognition Methods Integrating Knowledge Graph Constraints

By evaluating the link fingerprint and observation quality indicators of multimodal data on mobile devices and calibrating the edge weights of the knowledge graph edge by edge, the problems of signal quality fluctuation and individual differences in mobile depression assessment are solved, and the stability and low-cost operation and maintenance of depression screening and follow-up are achieved.

CN121565475BActive Publication Date: 2026-04-03SOUTH CHINA UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-26
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In mobile remote depression assessment, fluctuations in multimodal signal quality and individual differences lead to problems such as false alarms, biases, inconsistencies, and high maintenance costs in existing technologies for depression screening and follow-up. There is a lack of effective link calibration, conflict attribution, and edge-level adjustment mechanisms.

Method used

By acquiring multimodal data and calculating link fingerprints and observation quality indicators, the effectiveness of knowledge graph relationship edges is evaluated edge by edge. Edge weights and constraint strengths are adaptively calibrated to compensate for link-induced statistical morphological distortions and attribute conflict sources. An edge-level ledger is established for rapid rollback and recovery.

Benefits of technology

It improves the stability and interpretability of depression screening and follow-up, reduces false alarm rates and maintenance costs, and enhances cross-device consistency and individual adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565475B_ABST
    Figure CN121565475B_ABST
Patent Text Reader

Abstract

This application discloses a multimodal feature fusion and recognition method that integrates knowledge graph constraints, belonging to the field of computer data processing and multimodal intelligent recognition technology. It can be used for mobile remote depression screening, online psychological counseling, and intelligent follow-up assessment. The scheme involves: extracting audio, video, and text features and evaluating observation quality; processing link fingerprints on the end side based on signal statistical morphology and link statistical calculations, evaluating edge-by-edge and adaptively calibrating the strength of knowledge graph constraints; quantifying edge-level conflicts and attributing them to distinguish between observation degradation and relationship failure during fusion inference, adjusting modal contribution coefficients or edge constraints respectively, and forming an edge-level ledger to support ranking, localization, and rollback recovery. The technology of this application solves problems such as misjudgment of depression risk caused by link-induced false evidence, difficulty in attributing conflict sources, and difficulty in locating failed relationships, thereby improving the stability and maintainability of recognition results and risk trends under cross-device, cross-environment, and weak network conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer data processing and multimodal intelligent recognition technology, specifically to a multimodal feature fusion and recognition method that integrates knowledge graph constraints, applicable to application scenarios such as mobile remote psychological screening, online psychological counseling, and mental health follow-up. Background Technology

[0002] With the rapid development of mobile remote psychological screening, online psychological counseling, and intelligent follow-up applications, mental health assessment technologies based on multimodal data such as voice, video, and text are gradually becoming important tools for auxiliary screening and continuous monitoring of depression. By analyzing the speech rhythm and pause features, facial expressions and head movements / gaze features, emotional tendencies and cognitive cues in language content, and interaction behavior patterns generated by users during natural interactions, the system can, to a certain extent, reflect an individual's degree of depressed mood, signs of decreased interest, psychomotor retardation trends, and degree of negative expression, providing clinicians with follow-up references and risk warnings. In existing technologies, to improve the stability and interpretability of assessment results, some solutions introduce knowledge graphs or prior rules related to depression, such as the association between slower speech rate and retardation, the association between reduced facial expression activity and depressed mood, and the association between negative vocabulary and negative cognition, to guide or constrain the analysis process of multimodal features, thereby avoiding the uncertainty brought about by relying entirely on data-driven models.

[0003] However, in real-world applications, remote assessment data for depression typically originates from mobile terminals or remote interactive scenarios. This data is significantly affected by factors such as differences in device models, variations in end-user audio and video processing links, changes in environmental noise and lighting, network transmission fluctuations, and changes in usage habits and interaction states. Consequently, the quality and statistical characteristics of multimodal signals are prone to fluctuation. Furthermore, depression-related behaviors and expressions exhibit significant individual differences and stages of change, making fusion methods based on fixed prior relationships or static constraints prone to insufficient adaptability in long-term follow-ups. At the engineering level, these issues directly manifest as the sensitivity of depression risk scores to link status, fluctuating scores between sessions, and difficulty in interpreting and verifying trend curves. This can lead to false alarms causing unnecessary interventions or missed reports delaying follow-up appointments.

[0004] For example, Chinese invention patent with publication number CN113139062A discloses a depression detection system based on social media. It mainly preprocesses and quantifies social media text, constructs a depression knowledge graph, and then embeds it into a classification model. Finally, it uses sequence models such as LSTM to classify and train the text, so as to output auxiliary detection results of depression risk.

[0005] Existing technologies still have at least the following shortcomings:

[0006] First, the statistical distortion caused by the edge-side audio enhancement link leads to misjudgments in depression screening due to fixed graph edge weights or constraint strength. In engineering practice, mobile devices generally have processing links such as noise reduction, echo cancellation, automatic gain control, and encoding compression, and the processing strategies vary significantly among different brands and models of terminals, different system versions, and different call or recording modes. A common phenomenon is that some terminals weaken high-frequency components and suppress short-term energy fluctuations during noise reduction, making the speech sound flatter and more stable. As a result, prosodic features such as pitch variation amplitude, fundamental frequency jitter, energy fluctuation, and pause boundaries appear to have smaller numerical changes. At the same time, automatic gain control compresses the dynamic range, smoothing out the differences between loud and soft sounds. When there are relationships between small pitch variation and low mood, slow speech rate and lag in the knowledge graph, and the edge weights are fixed, the fusion model tends to output low mood or increased lag, thus raising the depression risk score. However, this tendency is not a true change in the depressive state, but rather caused by the edge-side processing link altering the signal spectrum and prosodic statistical morphology, which is a link-induced false evidence. Existing technologies mostly solidify or set the edge weights of the spectrum offline, lacking a mechanism to automatically calibrate the validity of the spectrum relationship by combining the end-side link status or proxy quantities such as frequency band energy distribution changes, dynamic range compression signs, and noise floor rise that can be calculated from voice data. This makes it difficult to avoid false alarms, biases, and inconsistencies when migrating across devices, thereby affecting the reliability of remote screening for depression in terms of risk change trends.

[0007] Secondly, the sources of constraint conflicts are difficult to attribute, and the confusion between relationship failure and observation degradation leads to opposite correction directions, causing fluctuations in the output of depression risk that are difficult to reproduce. Under conditions such as low light, side profile, occlusion, weak network, and frame loss, the confidence of video key points decreases, the clarity deteriorates, compression artifacts increase, and the frame rate fluctuates. As well as speech packet loss and transcription confidence fluctuations, these can lead to modal observation degradation, making certain modal features unusable or weakened. For example, under low light or occlusion conditions at night, face detection and key point localization are unstable, and facial micro-expressions and local muscle movements are difficult to capture stably, passively reducing features such as the intensity of facial expressions, eye contact, and head movement amplitude. If there is a relationship between reduced facial expressions and lag or low mood in the knowledge graph, conflicts are more likely to arise between graph inference and multimodal fusion evidence. Such conflicts often stem from observational degradation, such as poor visibility or unstable detection, rather than from the failure of the graph relations themselves. If graph constraints are directly weakened without source differentiation, the previously correct prior relations may remain unused for a long time even after illumination is restored or occlusion is removed, leading to insufficient utilization of valid evidence in follow-up assessments. Conversely, due to differences in dialectal speech rate, expression style, or individual baselines, the correspondence between certain speech rhythms, pause patterns, or facial expressions and depressive states may indeed weaken or reverse. If these are misjudged as short-term noise and the priors are continued to be applied forcefully, a systematic bias will form, causing persistently high or low depression risk scores and misleading follow-up judgments. Current technologies generally lack mechanisms for differentiating and treating conflict sources, making it difficult to determine whether to prioritize reducing modal contribution coefficients or graph edge constraint strength based on changes in observation quality. This results in fluctuating recognition output, poor stability, difficulty in interpretation, and difficulty in reproducing the results.

[0008] Third, the lack of edge-level traceable conflict quantification and operational closed-loop makes it difficult to locate failed relationships and achieve rapid rollback or recovery, resulting in high maintenance costs for depression follow-up systems. Knowledge graphs typically contain multiple relationship edges, and their failure sensitivity is highly correlated with device, environment, and network status. During online operation, the increase in false positives or false negatives of depression risk is often caused by a few key relationship edges being systematically triggered or continuously conflicted under specific link conditions. However, existing technologies mostly remain at the level of overall loss or overall weight, lacking the ability to quantify, accumulate, and sort which edge is continuously conflicting in the current scenario, and also lacking a traceable ledger that records changes in edge-level conflicts as the link status changes. A typical engineering phenomenon is that when network fluctuations lead to an increase in voice packet loss rate, a decrease in bitrate adaptation, or a reduction in transcription confidence, the overall false alarm rate increases. However, the lack of edge-level indicators makes it difficult to distinguish whether the conflict is a persistent conflict of speech prosody-related edges, a persistent conflict of text semantic-related edges, or a persistent conflict of video expression-related edges under specific lighting conditions. As a result, manual investigation, manual reconfiguration, or full retraining can only be relied upon, which is slow, costly, and prone to introducing new inconsistencies. It is difficult to meet the requirements of mobile remote depression screening and follow-up for continuous and stable operation and verifiable trend output.

[0009] In summary, existing technologies still struggle to achieve measurable evaluation, attributable adjustment, and closed-loop feedback of knowledge graph constraints under conditions of heterogeneous multi-device operation and fluctuating remote interaction links on mobile devices. There is an urgent need for a multimodal fusion and identification method for electronic digital data processing to suppress passive drift of graph relationships caused by edge links, weak network transcription, and environmental changes, and to improve the stability, interpretability, and maintainability of depression-assisted identification and long-term follow-up assessment. Summary of the Invention

[0010] In multi-device heterogeneous scenarios for remote mobile data acquisition, processing links such as edge-side noise reduction, echo cancellation, automatic gain control, and encoding compression, as well as transmission and transcription links such as packet loss in weak networks, bitrate adaptation, and transcription confidence fluctuations, can systematically alter the statistical form of audio and video and induce "link-induced false evidence," leading to the false triggering of fixed edge weights / static constraints in the knowledge graph. At the same time, under conditions such as low light, occlusion, and frame loss, observation degradation and relationship failure are easily confused, resulting in opposite correction directions and output fluctuations. Furthermore, the lack of edge-level conflict quantification and closed-loop ledger makes it difficult to locate failed relationships and quickly roll back and recover, resulting in high engineering maintenance costs.

[0011] To address the aforementioned technical problems, embodiments of the present invention provide a multimodal feature fusion and recognition method that integrates knowledge graph constraints. The technical solution is as follows:

[0012] A multimodal feature fusion and recognition method integrating knowledge graph constraints includes:

[0013] S1: Acquire multimodal data generated by the target object during natural interaction, perform feature extraction on each modality to obtain modal features for mental health risk assessment, and calculate observation quality indicators based on the multimodal data and objective information from its acquisition and transmission process. Multimodal data includes at least two or more of the following: audio data, video data, and text data. Observation quality indicators are used to characterize the reliability of the corresponding modality observations.

[0014] S2: Based on the statistical morphology and link statistics of multimodal data, calculate link fingerprints and obtain a knowledge graph related to mental health status identification. Link fingerprints characterize the systematic influence of end-side processing links and transmission transcription links on signal statistical properties. The knowledge graph contains relational edges associated with modal features.

[0015] S3: Based on the link fingerprint and observation quality indicators, evaluate the effectiveness of one or more relational edges one by one and adaptively calibrate the edge weights or constraint strengths of the corresponding relational edges.

[0016] S4: Perform multimodal feature fusion inference based on modal features and calibrated knowledge graph relation edges. The fusion inference assigns modal contribution coefficients to each modal feature, representing the weighted contribution ratio or gating ratio of each modal feature to the fusion output. During the fusion inference process, for one or more relation edges, edge-level conflict is calculated edge-by-edge to quantify the degree of inconsistency between the graph-constrained inference and the multimodal evidence. Conflict sources are attributed based on the edge-level conflict and observation quality indicators to distinguish between conflicts caused by observation degradation and conflicts caused by relation failure.

[0017] S5: Perform feedback adjustment based on the attribution results, where conflicts caused by observation degradation correspond to the adjustment of the modal contribution coefficient, and conflicts caused by relation failure correspond to the adjustment of relation edge weights or constraint strength.

[0018] S6: Based on the calibrated and feedback-adjusted fusion inference results, output the mental health status identification result. The identification result includes at least a risk score, risk level, or follow-up trend. The follow-up trend is the follow-up trend of the risk score and / or the follow-up trend of the stated risk level.

[0019] The beneficial effects of the technical solutions provided by the embodiments of the present invention include at least the following:

[0020] In mobile-based remote depression assisted identification and screening scenarios, a link fingerprint reflecting the influence of edge-side noise reduction, automatic gain control, and encoding compression is calculated. Based on this fingerprint, the constraint strength of edges related to prosodic changes and speech rate pauses in the knowledge graph is calibrated edge-by-edge. This achieves automatic compensation for link-induced statistical morphological distortions, preventing changes caused by edge-side processing, such as smaller intonation fluctuations and smaller energy fluctuations, from being directly interpreted as evidence of increased depression risk. Compared to existing technologies that solidify or set graph edge weights offline, which are prone to misjudging link effects as low mood and sluggishness when migrating across different device models and system versions, this approach reduces false positives in depression screening and improves cross-device consistency and scoring stability.

[0021] In the scenario of depression identification during remote consultation and follow-up in weak network conditions, this method calculates edge-level conflict quantities on relational edges and combines them with observational quality indicators such as transcription confidence, latency jitter, and packet loss rate to attribute the source of conflicts. This allows for the differentiation and processing of unreliable semantic evidence caused by transcription link fluctuations from changes in actual language content, avoiding sudden increases in depression risk scores due to transcription errors such as the loss of negative words or the replacement of emotional words. Compared to existing technologies that uniformly weaken graph constraints or uniformly reduce modal weights, making it difficult to explain the reasons for fluctuating scores in weak network conditions, this method improves the reproducibility and interpretability of depression risk assessment under weak network conditions.

[0022] In depression identification scenarios under degraded mobile acquisition conditions such as low light occlusion and frame drops, this method attributes edge-level conflicts based on indicators such as keypoint confidence, effective frame ratio, and clarity. When observations degrade, it prioritizes lowering the contribution coefficient of facial expression-related modalities rather than directly weakening prior relation edges. This avoids the passive decrease in facial expression intensity due to poor visibility being mistakenly interpreted as evidence of diminished interest or psychomotor retardation. Compared to existing technologies that directly weaken graph constraints or continue to forcefully apply priors when video is unstable, which easily leads to jitter in depression risk output and is difficult to verify, this method improves the stable capture capability of retardation-related cues in remote video follow-up.

[0023] In scenarios requiring continuous and consistent assessment for long-term follow-up of depression, this approach establishes a boundary-level ledger to record link fingerprints, boundary-level conflict quantities, and adjustment actions, and performs cumulative sorting and localization. Simultaneously, it sets limits, hysteresis, and cooling conditions for rollback and recovery. This enables rapid localization and smooth recovery of a few key relationship edges that are continuously mistriggered under specific link conditions, preventing prolonged elevation of depression risk scores even after short-term network weakness or temporary occlusion has ended. Compared to existing technologies that rely on overall loss or overall weight, and where online anomalies often require manual investigation or full retraining with difficulty in rollback, this approach reduces the operational costs of the follow-up system and improves the long-term consistency of risk scores.

[0024] In remote follow-up scenarios where significant individual differences exist in depression identification, this method establishes individual baselines and adaptively adjusts conflict levels and failure thresholds. It aligns individuals with long-term stable characteristics such as a steady tone of voice, minimal facial expressions, and slow speech rate as their personal baselines, reducing the probability of misclassification as depressed mood or lethargy. Compared to existing technologies that use uniform group thresholds and are prone to systematic false positives for specific expression styles, this method improves the fairness and cross-population generalization of depression risk assessment, and makes follow-up trends more consistent with actual individual changes.

[0025] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0026] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this application. Wherein:

[0027] Figure 1 This is a flowchart illustrating a multimodal feature fusion and recognition method that integrates knowledge graph constraints, as provided in an embodiment of the present invention.

[0028] Figure 2 This is the end-to-network-to-cloud collaborative architecture provided in the embodiments of the present invention;

[0029] Figure 3This is a schematic diagram of a sub-graph of the knowledge graph related to depression provided in an embodiment of the present invention;

[0030] Figure 4 This is a diagram of the edge mechanism provided in an embodiment of the present invention. Detailed Implementation

[0031] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details of the embodiments to aid understanding, and should be considered merely exemplary.

[0032] like Figure 2 As shown, this embodiment adopts an end-network-cloud collaborative architecture: the end side completes multimodal acquisition, feature extraction and observation quality / link fingerprint statistics and reporting, the network side provides transmission statistics, and the cloud side provides transcription, knowledge graph side tables and fusion reasoning services, and forms a closed loop regulation through ledger and configuration distribution.

[0033] Embodiment 1 of the present invention: Figure 1 As shown, Figure 1 This is a flowchart illustrating a multimodal feature fusion and recognition method incorporating knowledge graph constraints, provided by an embodiment of the present invention. This embodiment offers a multimodal feature fusion and recognition method incorporating knowledge graph constraints, applicable to scenarios such as mobile remote mental health status assisted screening, online psychological counseling, and intelligent follow-up assessment. The specific steps are as follows:

[0034] S1: The mobile device collects multimodal data of the target object during a natural interaction session. Since any modality may become unavailable due to permissions, environment, or network degradation in mobile remote interaction scenarios, the multimodal data used includes at least two or more of the following: audio data, video data, and text data. The multimodal data includes the raw data of the corresponding modality and its acquired metadata, and further includes objective quality and link information returned by the system interface or calculated statistically from the raw data. The raw data is used for subsequent feature extraction, the acquired metadata is used for temporal alignment and resampling processing, and the objective quality and link information is used for calculating observation quality indicators and link fingerprints. To facilitate cross-modal fusion and link impact analysis, the collected metadata includes at least the timestamp sequence of each modality, sampling rate / frame rate configuration, resolution, audio channel mode, encoding mode identifier, and session round time stamp. The timestamp is recorded by the media frame or acquisition thread at the acquisition time and carried with the data frame. The system uses the timestamp as a unified benchmark to align audio frames, video frames, and text segments. When there are fluctuations in sampling rate or frame rate, resampling or interpolation is used to aggregate the feature statistics into a unified statistical time window. At the same time, the alignment error and missing ratio are used as objective inputs for subsequent observation quality assessment.

[0035] The audio data includes at least an audio waveform frame sequence and its timestamp sequence. The audio waveform frame sequence is an audio stream acquired by the terminal microphone. Based on the audio waveform frame sequence, objective statistics such as the noise floor of non-speech segments, short-time energy sequence, and the proportion of effective speech segments can be further obtained. On the audio side, the short-time energy sequence is calculated by dividing the audio waveform into frames according to the frame length. The short-time energy of each frame can be obtained by the mean or sum of the squares of the sample values ​​in that frame. The proportion of effective speech segments is obtained by speech activity detection. Speech activity detection can be obtained by the speech / non-speech markers output by the terminal speech activity detection model, or by the speech segment markers returned by the call / recording system interface. Based on this, the proportion of speech segment duration to the total session duration is calculated. The noise floor of non-speech segments is calculated from the non-speech segment samples obtained by speech activity detection. At least the low quantile or mean of the short-time energy of non-speech segments can be used as the noise floor statistic to characterize the level of environmental noise and residual noise from terminal noise reduction. The video data includes at least a video frame sequence and its timestamp for each frame. The video frame sequence is a video frame stream captured by the front-facing camera. The frame rate and resolution are returned by the camera configuration. If there are frame rate fluctuations, the actual frame rate sequence is obtained by statistically analyzing the differences in the video frame timestamps. Furthermore, based on the video frame sequence, objective quantities such as face detection results, key point coordinates and key point confidence, effective frame ratio, sharpness, or low light level can be obtained. Among them, the face detection results and key point confidence are output by the visual model and can be directly statistically analyzed. The sharpness or low light level is calculated by statistically analyzing the image brightness distribution and gradient. On the video side, face detection results, keypoint coordinates, and keypoint confidence scores are obtained from the visual model's inference output for each frame of video image. The keypoint confidence score is the model's output value of the reliability of each keypoint location and is returned with each frame. Valid frames are those that meet preset usability conditions. Usability conditions include at least successful face detection and keypoint confidence scores reaching the minimum usability condition. The minimum usability condition can be generated by quantile thresholds obtained from historical stable session samples, or dynamically updated and written to the configuration table by online rolling quantile statistics. The percentage of valid frames is obtained by the ratio of the number of valid frames to the total number of frames. Sharpness is calculated statistically from image gradients, and can be characterized by calculating the gradient magnitude of image frames and statistically analyzing their mean or variance. Low-light intensity is calculated statistically from image brightness distribution, and can be obtained by statistically analyzing the brightness mean, low-light pixel proportion, or brightness quantile of image frames. These, along with sharpness statistics, serve as objective inputs for video observation quality indicators. The text data includes at least the user input text and its input timestamp, as well as the speech-transcribed text and its corresponding timestamp. The speech-transcribed text is generated and returned by a local or cloud-based speech recognition service, which also returns the confidence or stability score of each transcription to characterize the reliability of the text modality. Furthermore, it can further collect interactive metadata such as the number of words in each round of dialogue, the round structure, and the response latency to help describe the changes in the natural interaction state.On the text side, the speech-to-text is generated and returned by a local or cloud-based speech recognition service. The transcription service can return the confidence or stability score for each transcribed segment. When the transcription service does not directly provide a confidence field, stability indicators, candidate transcription consistency statistics, or post-decoding scores within the speech recognition service can be used as alternative confidence levels. The source type of the confidence level is marked in the conversation record to ensure the traceability of subsequent fusion inference and conflict attribution. The number of words, round structure, and response latency in each round of dialogue are recorded and statistically analyzed by the interaction system when input and return events occur. In addition, objective information from the acquisition and transmission process is obtained through the system network statistics interface or transport layer statistics, including at least packet loss rate, latency, jitter, retransmission count, or adaptive bitrate changes. This objective information is directly read from the network stack or transport layer statistics and summarized by time window. By acquiring at least two modalities, usable evidence can be provided by another modality when a single modality degenerates, and the edge-level conflict attribution and feedback adjustment mechanism between knowledge graph constraints and multimodal evidence can be maintained. Packet loss rate, latency, jitter, retransmission count, and adaptive bitrate change can be obtained from the system network stack statistics interface, transport layer statistics module, or media transmission module's real-time statistics interface. Alternatively, they can be obtained from the application layer's sending and receiving logs and summarized according to a statistical time window consistent with audio and video. Among these, packet loss rate is at least obtained from the sent packet count and lost packet count; latency is at least obtained from the end-to-end arrival time difference or round-trip time delay; jitter is at least obtained from the fluctuation of adjacent packet arrival intervals or latency sequences; retransmission count is obtained from the transport layer's retransmission event statistics; and adaptive bitrate change is obtained from the encoder or transport layer's bitrate adjustment log statistics. The above link statistics, aligned with the session timestamp, are used as link information input for subsequent link fingerprint calculation.

[0036] After acquiring the data, modal features were extracted for each modality: Audio features were calculated by segmenting the audio stream into frames, including speech rate-related features, pause ratio, fundamental frequency-related features, and short-time energy fluctuation-related features; the fundamental frequency was calculated naturally from the audio periodicity, and the short-time energy was calculated from the sum of squares of each frame. Video features were obtained by performing face detection and keypoint localization on video frames, including facial expression intensity, facial motion amplitude, and gaze stability; keypoint confidence was directly output by the keypoint localization model and returned with each frame. Text features were calculated by segmenting or encoding the text, including sentiment tendency, intensity of negative expression, and proportion of negative expression; the semantic confidence of the text could be obtained from the text sentiment model output, or the transcription confidence could be used as the main source of text reliability. Simultaneously, observation quality indicators were calculated based on the objective information of the multimodal data and its acquisition and transmission processes, including audio observation quality indicators, video observation quality indicators, text observation quality indicators, and transmission observation quality indicators. Audio observation quality metrics are calculated statistically from audio signals, including signal-to-noise ratio estimation, noise floor in non-speech segments, and the proportion of effective speech segments; effective speech segments are calculated from speech activity detection. Video observation quality metrics are obtained statistically from video frame processing and model output, including the proportion of effective frames, average confidence of key points, and sharpness statistics; effective frames are those with successful face detection and key point confidence exceeding the minimum usable threshold, and sharpness is calculated statistically from image gradients. Text observation quality metrics are returned and statistically obtained from the transcription service, including the mean transcription confidence and fluctuation range. Transmission observation quality metrics are directly obtained from the system network statistics interface, including packet loss rate, latency, and jitter, and summarized by session time window. Processing parameters such as audio frame length and statistical time window length are engineering implementation parameters, and their initial ranges can be given by historical stability statistics. The minimum availability conditions, normalization intervals, and thresholds and quantile boundaries of each quality statistic can be generated from historical stable sample statistics as cold start parameters, and updated using online rolling quantile statistics during online operation; the generation time window, version number, and update strategy of each parameter can be recorded in the configuration table or model version information.

[0037] S2: Statistical Morphology and Link Statistical Calculation of Link Fingerprint Based on Multimodal Data. Link fingerprint is used to characterize the systematic influence of end-side processing links and transmission transcription links on signal statistical characteristics. Statistical morphology inputs come from computable statistics inherent to audio and video: audio short-time energy distribution, frequency band energy ratio, non-speech segment noise floor, video frame interval fluctuations, and compression artifact approximation strength, etc. On the audio side, the audio short-time energy distribution is obtained by statistically analyzing the short-time energy sequence obtained from segmenting the audio waveform. At least the high and low quantile differences, kurtosis, or energy fluctuation amplitude of the short-time energy sequence can be used to characterize whether energy fluctuations are compressed. The frequency band energy ratio is obtained by performing frequency band filtering or short-time spectrum analysis on audio frames. At least the ratio or relative proportion change of low-frequency energy to high-frequency energy can be statistically analyzed. The non-speech segment noise floor is obtained by statistically analyzing non-speech segment samples obtained from speech activity detection. The low quantile or mean of the non-speech segment short-time energy can be used as the noise floor statistic. On the video side, the fluctuation of video frame interval is obtained by statistical analysis of the frame interval sequence obtained by the difference of timestamps of adjacent video frames. At least the standard deviation, quantile difference, or out-of-limit ratio of the frame interval sequence can be used to characterize the frame rate jitter. The approximate intensity of compression artifacts is obtained by image statistics. At least the degree of high-frequency energy attenuation of the image, the proportion of gradient anomalies at block boundaries, or the degree of texture consistency reduction can be used as surrogate statistics. All of the above statistics are objective calculable quantities of the original audio and video data and can be repeatedly calculated within the time window.

[0038] Link statistics inputs come from system and service returns: encoding parameters, bitrate adaptive changes, packet loss rate, transcription confidence fluctuations, etc., which are obtained and summarized from media encoding module logs, network statistics interfaces, and transcription service return values, respectively. To adapt to differences in different device models and network environments, the normalized intervals of each component of the link fingerprint can be generated from quantile boundaries obtained from historical stable session samples. When online data is insufficient, an initial interval can be provided using a cold start configuration table, and the version number can be recorded. Encoding parameters and bitrate adaptive changes can be obtained from the operation logs and statistics interfaces of the media encoding module or real-time audio and video transmission module, including at least obtainable fields such as target bitrate, actual bitrate, keyframe interval, quantization intensity, or encoding level changes. Based on the time series of the fields within the statistical time window, the change amplitude or fluctuation degree of the corresponding fields is calculated. Packet loss rate, latency, and jitter are obtained from the system network stack statistics interface or transport layer statistics module and summarized according to the statistical time window. Transcription confidence fluctuations are obtained from the confidence sequence returned by the speech recognition service, and at least the mean, quantile difference, or fluctuation amplitude can be used to characterize transcription stability. The above link statistics are aligned with audio and video timestamps when summarized, so that the link fingerprint can reflect the synchronous impact of link changes on signal morphology.

[0039] Simultaneously, a knowledge graph related to mental health status identification is acquired. The knowledge graph contains relational edges associated with modal features, which express the constraints between features and states. The knowledge graph can be pre-built and stored in a graph database or stored on a server as a relational edge table. At the start of a session, the mobile device retrieves the graph version number and caches relational edges and their weights or constraint strengths as needed. To support edge-by-edge evaluation, edge-by-edge calibration, and edge-level conflict localization, relational edges can be organized in the form of relational edge tables during storage. These tables include at least the relational edge identifier, relational edge type, associated modal type, edge weight or constraint strength, graph version number, and the identifier of the feature item the edge depends on. The relational edge type distinguishes different associations such as prosody-emotion, facial expression-lag, and negative semantics-negative cognition. The associated modal type indicates that the edge primarily relies on audio, video, text, or cross-modal consistency evidence, facilitating the selection of corresponding observation quality indicators and link fingerprint components in S3 for edge-by-edge effectiveness evaluation. The initial relation edges and initial edge weights of a knowledge graph can come from sources such as: the strength of association between features and risk labels obtained from historical labeled samples, where the statistics can be calculated using objective statistics such as correlation or mutual information; structural relationships established based on the mapping between clinical scale dimensions and symptom dimensions, with edge weights then calibrated using historical data; and when historical data is insufficient, edge weights can be initialized with uniform initial values ​​and recorded as graph version parameters, subsequently replaced gradually by online calibration. Figure 3 As shown, in another implementation, mental health status identification is depression risk identification, and the knowledge graph is a knowledge graph related to depression.

[0040] To ensure a clear and easily verifiable output structure for the link fingerprint, this embodiment represents the link fingerprint as a vector or structure, which includes at least one or more of the following: audio morphology components, video temporal components, video compression components, network transmission components, and transcription stability components. Each component is statistically obtained within a sliding time window and aligned with a unified session timestamp. For example, audio morphology components may include short-time energy quantile variation, band power ratio variation, and noise floor rise; video temporal components may include frame interval fluctuations and frame rate fluctuations; video compression components may include the approximate intensity of compression artifacts; network transmission components may include packet loss rate, mean latency, and jitter; and transcription stability components may include confidence quantile and the proportion of low-confidence segments. These components can be selectively enabled based on device availability, and component availability flags are written into the session record to prevent the link fingerprint from being unavailable due to a single interface. To ensure cross-device comparability, each component can be truncated and normalized using the rolling quantile boundaries of historically stable samples or recently stable samples: for any component x, take the lower bound L_x and the upper bound U_x, and map them as... Where x_norm represents the dimensionless normalized value after interval normalization of component x. The function defined by the brackets restricts the calculation result within the range of 0-1: 0 for a result less than 0, 1 for a result greater than 1, and the result itself in all other cases. L_x and U_x are generated by rolling quantile statistics and recorded with each version. All components of the normalized link fingerprint fall within the 0-1 range; larger values ​​indicate a stronger systematic influence of the link on the statistical morphology. The compressed artifact approximation intensity provides a recalculated calculation method: the system divides each frame of image into several blocks of a preset block size, selects boundary zone pixels near the block boundaries and calculates the boundary gradient mean G_b, while simultaneously calculating the internal gradient mean G_in within the block region to obtain the block effect ratio. The calculation uses the percentage of blocks exceeding a threshold (R_blk) as a block effect metric. ε_0 is a positive numerical stability constant to prevent computational instability caused by G_in approaching zero. ε_0 can be determined by the low quantile value of G_in obtained from historical stable samples, or dynamically updated by online rolling quantile statistics, and recorded in the configuration table. Simultaneously, high-pass filtering or short-time spectrum / discrete cosine transform can be performed on the image to calculate the proportion of high-frequency energy and its relative attenuation as a detail loss metric. The compression artifact approximation intensity can be composed of the block effect metric and the detail loss metric within a time window, for example, taking one of them or their joint summary. ε_0 and the threshold are written into the configuration and the version number is recorded. This calculation does not rely on vendor-specific encoding parameter interfaces and can still be objectively calculated from the original video frames when encoding logs cannot be read, thus ensuring the availability of the link fingerprint.

[0041] S3: Evaluate the effectiveness of each edge and adaptively calibrate the edge weights or constraint strength. Based on link fingerprints and observation quality indicators, evaluate the effectiveness of at least a portion of the relation edges one by one, and adaptively calibrate the edge weights or constraint strength of the corresponding relation edges. Each relation edge in the knowledge graph is associated with and records at least the following information in the edge table: a relation edge identifier for uniquely locating the edge, a relation type for indicating the symptom dimension or behavioral cue type of the constraint of the edge, one or more major dependency modalities among audio, video, text, or transcription, and evidence binding rules that can be used to indicate which modal features the evidence corresponding to the edge is mapped from. Specifically, the system records at least the following for each relation edge in the edge table: the edge mainly depends on the modality type, the list of evidence features, and the aggregation method; the list of evidence features is selected from the modality features extracted by S1, such as speech rate / pause ratio / fundamental frequency fluctuation, facial expression intensity / head movement gaze stability, negative word ratio / negative expression ratio, etc.; the aggregation method is a recalculated statistical aggregation method, such as normalizing the evidence features and taking a weighted average or mean to obtain the evidence quantity of the edge within the time window, and using the evidence quantity for subsequent inference difference calculation and edge-level conflict assessment. Examples include speech rate / pauses, fundamental frequency and energy fluctuations, facial expression intensity and head movement / gazing stability, negative word ratio and negation expression ratio, and link fingerprint binding rules, which specify which type of link fingerprint component should be used when evaluating the effectiveness of an edge. Examples include audio morphological distortion fingerprints caused by edge-side noise reduction / automatic gain / encoding compression, video link fingerprints caused by frame rate fluctuations / compression artifacts, and transcription link fingerprints caused by transcription confidence fluctuations and packet loss jitter. The initial edge weight or initial constraint strength of the edge is stored as a graph version parameter with the version number. All observation quality indicators come from S1: such as audio signal-to-noise ratio estimation, effective speech segment ratio, effective video frame ratio, keypoint confidence, sharpness statistics, transcription confidence statistics, packet loss rate / latency jitter statistics, etc. All link fingerprints come from S2: such as short-time energy quantile difference changes, frequency band energy ratio changes, frame interval fluctuations, compression artifact approximation strength, encoding parameter changes, bitrate adaptive changes, transcription confidence fluctuations, etc. All of the above are obtained through statistical calculations of raw data or summarization of system interface / service return values.

[0042] During edge-by-edge evaluation, for each relation edge, a corresponding observation quality score is first selected based on its primary dependent modality type. This quality score is then normalized to the 0–1 range using rolling quantile statistics of historical stable samples or online stable samples, with scores closer to 1 indicating more reliable modality observations. Simultaneously, a corresponding link fingerprint component is selected based on its link fingerprint binding rule, and similarly normalized to the 0–1 range using historical statistics or rolling quantile statistics, with scores closer to 1 indicating a stronger systematic influence of the link on the statistical pattern and a greater likelihood of inducing link-induced false evidence. Based on this, the edge-level validity coefficient within the current session / time window is calculated to characterize the credibility of the edge for constrained fusion inference under the current link and quality conditions. This coefficient ranges from 0 to 1 and satisfies the following conditions: the coefficient continuously decreases as observation quality decreases or link influence strengthens, and continuously increases as observation quality recovers and link influence weakens. This embodiment allows for the following reproducible calculation method that does not rely on subjective thresholds: multiplicative suppression is applied to the "normalized value of observation quality" and the "normalized value of link impact", that is, the edge-level effectiveness coefficient can be obtained by multiplying the "normalized value of observation quality" by "(1 minus the normalized value of link impact)", and the amplitude is limited within the range of 0-1; the upper and lower bounds of normalization are maintained by rolling quantile statistics and the version number is recorded for verification.

[0043] After calculating the edge-level effectiveness coefficient, a limited-amplitude smooth update is performed on the edge weight or constraint strength of the edge: first, the target constraint strength of the edge is obtained based on the edge-level effectiveness coefficient, and the target constraint strength is made to change continuously with the effectiveness coefficient; then, the current edge weight or constraint strength is smoothly brought closer to the target constraint strength, while setting upper and lower limits and a maximum single update amplitude to avoid discrete jumps caused by short-term fluctuations. The smooth update can be implemented using exponential smoothing or sliding window smoothing; the initial values ​​of the maximum single update amplitude and upper and lower limits can be determined from offline historical data through a joint objective of reducing the false judgment rate and reducing output fluctuations; when there is insufficient online data, a smaller update amplitude can be used and gradually increased; after going online, the update rate can be automatically fine-tuned based on the fluctuation statistics of stable sessions and the configuration version can be recorded, thereby ensuring that the edge weight changes are stable and traceable.

[0044] S4: As Figure 4As shown, in this embodiment, during the fusion inference process, both the evidence quantity and the inference quantity are obtained for each relation edge. The difference between the two is calculated to form the edge-level conflict quantity, and the conflict is attributed. When the attribution is due to observation degradation, short-term adjustment is performed to suppress the influence of short-term noise. When the attribution is due to relation failure, long-term adjustment is performed to smoothly adjust the edge weights or constraint strength. At the same time, the attribution and adjustment results are written into the edge-level ledger and support rollback recovery. Based on modal features and calibrated knowledge graph relation edges, multimodal feature fusion inference is performed. Fusion inference can be implemented by gated weighted fusion or fusion scoring with graph constraint terms. During the fusion process, the fusion results on the multimodal evidence side and the inference results on the graph constraint side are output. In an optional implementation, using a unified statistical time window consistent with S1 as the unit, the evidence scores of the three modalities of audio, video, and text are calculated for each time window t, and gating weights are generated based on the observation quality scores of each modality to achieve gated weighted fusion. Specifically: for each modality m∈{a(audio), v(video), x(text)}, the modal features of that modality within the time window t are denoted as... And through a recomputable scoring mapping function. Obtain the evidence score for this modality. ,in ∈[0,1]; rating mapping function This can be achieved by linear mapping followed by Sigmoid normalization, or by taking the risk probability output from the end-side model and directly normalizing it to the 0–1 interval. The modal observation quality score obtained from S1 is denoted as... ∈[0,1], generate gating weights For example, using normalized weighting: Where j∈M indicates that j is any modality index in set M, M is the currently available modality set, and ε is a preset small positive stable term used to avoid a denominator of zero. The gated multimodal evidence fusion score is obtained as follows: Therefore, when a certain mode is subjected to weak mesh, low light, occlusion, or unstable transcription conditions, it may lead to... As the threshold decreases, the gating weights will decrease continuously, thereby causing the fused output E_fuse(t) to converge continuously and remain stable.

[0045] This embodiment calculates edge-level conflict quantity for at least some relation edges on an edge-by-edge basis. The calculation of edge-level conflict quantity includes at least the following two reproducible objects: the evidence quantity of the edge, which is obtained by mapping the evidence features bound to the edge. The mapping rules are stored with the graph version and can be verified. For example, edges with slower speech rate map speech rate and pause ratio to lag evidence strength; edges with reduced facial expression activity map facial expression activity intensity and facial movement amplitude to low activity evidence strength; and edges with negative semantics map the proportion of negative words and the proportion of negative expressions to negative cognitive evidence strength. The inference quantity of the edge is formed by the relation type, direction, and calibrated edge weight or constraint strength of the edge. In this embodiment, edge-level inference quantity and edge-level constraint deviation are defined for each relation edge z within a time window t, and are used as graph constraint items in the fusion scoring. Specifically: for relation edge z, the edge table records at least its starting node u(z), ending node v(z), relation type type(z), direction dir(z), and edge weight or constraint strength after S3 calibration. The relation type is used to specify the sign and direction of the edge. When type(z) represents a positive correlation / promotion / co-directional constraint, take the following value: =1; when typz(z) represents negative correlation / suppression / reverse constraint, take 1. =–1; When type(z) is of other types, it can be explicitly stored in the edge table. Values ​​are taken for verification. The current state evidence value corresponding to node u(z) is denoted as S_u(t)∈[0,1], and the current state evidence value corresponding to node v(z) is denoted as... Where S_u(t) and S_v(t) can be obtained from the gated fused multimodal evidence mapping, or from the evidence feature mapping bound to the edge and normalized to the 0–1 interval. The edge-level inference quantity of this edge within the time window t is defined as: ;where clip(·) is used to restrict the value to the range of 0–1. Inference quantity By relation type (type→ ), direction (dir is determined by u→v) and calibrated edge weights Together they form. The edge-level constraint deviation within the time window t is defined as follows (which can also be considered a simplified definition of edge-level conflict): The constraint deviations of all relation edges are summarized to form graph constraint terms, for example: ,or , where Z is the set of enabled relation edges. Fusion score with graph constraints: λ represents the constraint weight, which can be recorded as a configuration parameter and managed with the version number. The output R(t) of the fusion inference is simultaneously affected by the multimodal evidence fusion score and the graph constraint term, and all quantities can be recalculated from the session data and edge table parameters. The edge-level conflict quantity is obtained by normalizing the difference between the evidence quantity and the inference quantity, and the introduction of this edge during the normalization process mainly depends on the observation quality score corresponding to the modality: when the observation quality decreases, the conflict quantity will be suppressed to reduce the amplification of the conflict quantity by false inconsistencies introduced by noise, occlusion, weak network frame loss, or transcription instability; when the observation quality remains in the normal range but the difference still exists, the conflict quantity remains high. At the same time, a minimum positive constant is added to the normalization denominator. This constant can be taken as the low quantile value of the observation quality score in the stable samples or preset as a small constant, and is recorded with the version number. Subsequently, conflict sources are attributed based on the amount of edge-level conflict and the observation quality score to distinguish between conflicts caused by observation degradation and conflicts caused by relationship failure: when the increase in edge-level conflict and the decrease in observation quality score occur simultaneously within a time window, and the observation quality score is lower than the usable lower bound obtained from the rolling quantile statistics of the upper stable samples, the conflict is attributed to observation degradation; when the observation quality score remains within the normal range (also maintained by the rolling quantile statistics) while the amount of edge-level conflict remains high within a continuous time window or accumulates to exceed the persistence threshold, the conflict is attributed to relationship failure. The normal range threshold and persistence threshold are generated by historical statistics or the upper rolling quantile statistics and are updated periodically to avoid fixed manual settings; the length of the continuous time window and the cumulative scope can be achieved by exceeding the threshold for several consecutive time windows or by exceeding the threshold percentage within a sliding window, and the parameters and version number are written into the configuration for verification.

[0046] S5: Execute feedback adjustment based on the attribution results. Conflicts caused by observation degradation correspond to adjustments in the modality contribution coefficient, while conflicts caused by relation failure correspond to adjustments in relation edge weights or constraint strength. To avoid short-term degradation from harming prior relations, this embodiment combines short-term and long-term feedback adjustment. When a conflict is attributed to observation degradation, short-term feedback adjustment is triggered: prioritizing a reduction in the contribution of the corresponding degraded modality to the fusion output, making the fusion inference output more stable under short-term degradation conditions such as weak networks, low light, occlusion, frame drops, or unstable transcription. The update amount of the modality contribution coefficient is determined by the decrease in the modality's observation quality score and the amount of edge-level conflicts associated with it. The worse the observation quality and the more significant the conflict, the lower the modality contribution coefficient. Simultaneously, a lower bound for the contribution coefficient is set to avoid completely removing the modality, which would lead to information breakage, and a smoothing method is used for contribution coefficient updates. The lower bound for the contribution coefficient can be determined by the lowest available contribution coefficient obtained from historical data statistics, or dynamically updated and recorded by online stable samples. When conflicts arise due to relationship failure, long-term feedback adjustment is triggered: only when the cumulative evaluation of edge-level conflict across time windows meets the persistence condition is the edge weight or constraint strength of the corresponding relationship edge weakened in small steps to avoid long-term weakening of the relationship edge due to short-term observation degradation. The triggering conditions for long-term feedback adjustment include at least: the corresponding modal observation quality score remains within the normal range within the continuous time window and the edge-level conflict amount meets the persistence threshold; if the observation quality score is unstable or below the usable lower bound, the long-term weakening of the edge is suspended, and only the short-term suppression of the modal contribution coefficient by short-term feedback adjustment is retained. Long-term feedback adjustment also adopts amplitude-limited smoothing update and sets a cooling and hysteresis mechanism: once a weakening update is triggered for an edge, it will not be triggered again within the cooling time window; when the subsequent observation quality recovers and the edge-level conflict amount is below the recovery threshold within the continuous time window, the edge weight or constraint strength of the edge is allowed to smoothly return to a more reasonable range. The cooling time window, recovery threshold, and recovery conditions can all be generated by historical statistics or online rolling quantile statistics and recorded together with the map version number and threshold version number for traceability. After completing the above adjustments, this embodiment can write the edge identifier, current link fingerprint summary, observation quality score, edge-level conflict amount, attribution result, and triggered adjustment action of each relationship edge that has undergone calibration or adjustment into the edge-level ledger, providing a basis for subsequent sorting, location, verification, and rollback recovery.

[0047] S6: Output the mental health status identification result based on the fused inference result after calibration and feedback adjustment. The identification result includes at least a risk score, risk level, or its follow-up trend. In this embodiment, the risk score is jointly formed by the multimodal evidence-side fusion result and the graph constraint-side inference result: the multimodal evidence-side fusion result has adaptively adjusted the contribution coefficients of each modality based on the observation quality score under the short-term feedback adjustment in S5, and the graph constraint-side inference result has adaptively adjusted the edge weights or constraint strengths of each relation edge under the edge-by-edge calibration in S3 and the long-term feedback adjustment in S5, thereby making the risk score more robust to fluctuations across devices, environments, and weak networks. The risk level threshold can be determined by the statistical distribution of historical labeled samples, or dynamically updated and recorded by the rolling quantile statistics of online stable samples to adapt to changes in population and device distribution. The follow-up trend is generated from a sequence of risk scores for consecutive sessions, and quality-weighted smoothing is applied based on the observation quality score of each session. When a session is in a weak network, low light, occlusion, or unstable transcription leading to low observation quality, the weight of that session on the trend is reduced accordingly, thereby minimizing the interference of degraded sessions on trend judgment. When session quality recovers and conflict levels decrease, the trend output can smoothly regress and reflect a more reliable direction of risk change. Quality-weighted smoothing can be implemented using a sliding window weighted average or exponential smoothing, with weights directly generated from the observation quality scores. The smoothing window length or smoothing coefficient is recorded and configured as an engineering parameter to ensure that the trend generation process is reproducible, interpretable, and verifiable. In this embodiment of the invention, let the risk score of the q-th session be... The observation quality summary is The trend can then be generated exponentially smoothed: Where ρ is the smoothing coefficient and the version number is recorded; when When the risk level is low, the session automatically reduces the magnitude of trend updates, thereby suppressing the interference of degraded sessions on trend judgment. This includes the session risk score. To quantify the psychological health status risk of the system's output for the q-th session, the risk output value obtained through fusion inference can be taken and normalized to the range of 0 to 1 using a unified dimension. The larger the value, the higher the risk level. The observation quality summary is obtained by summarizing the audio observation quality score, video observation quality score, text observation quality score, and transmission observation quality score of this session. During the summarization, only the available modalities are aggregated, and the results are normalized to the range of 0 to 1. The larger the value, the more reliable the observation of this session. The audio observation quality score is obtained by weighted summation of signal-to-noise ratio estimation, effective speech segment ratio, and noise floor statistics after normalization within a statistical time window; the video observation quality score is obtained by weighted summation of effective frame ratio, average confidence of key points, and sharpness statistics after normalization; the text observation quality score is obtained by summation of the mean transcription confidence and its fluctuation amplitude after normalization; the transmission observation quality score is obtained by normalizing packet loss rate, latency, and jitter, and then calculating and summing them according to the direction of lower packet loss rate / latency / jitter, which indicates higher quality. The upper and lower bounds of the above normalization are generated by statistical analysis of historical stable samples or online rolling quantiles, and the version number is recorded for verification. In some implementations, the mental health status identification result can be the depression risk identification result. For example... Figure 2 As shown, the identification results at this time include a depression risk score, a depression risk level mapped from the depression risk score, and the trend of change of the depression risk score and / or the depression risk level within the follow-up time window. Accordingly, the knowledge graph can be a depression-related knowledge graph, or a subgraph extracted from a depression-related knowledge graph corresponding to the modal features of the current conversation.

[0048] Embodiment 2 of the present invention: Based on Embodiment 1, this embodiment further illustrates the specific acquisition method of observation quality indicators and link fingerprints, so that the system has a consistent data acquisition caliber for measuring the quality level and the strength of the link's impact under different terminal models, different environmental noise and different network conditions, thereby providing a verifiable basis for subsequent edge-by-edge calibration, edge-level conflict quantity calculation and conflict source attribution.

[0049] To ensure the stability of statistical results and facilitate cross-session comparison, in this embodiment, all quality indicators and link fingerprints are calculated and updated according to the intra-session sliding time window. The length of the sliding time window can be determined by the principle of minimizing indicator fluctuations and responding promptly to state changes based on historical stable samples, and fine-tuning is allowed after going live based on the fluctuation statistics of stable samples. The time window length, update frequency, and statistical scope are recorded as configuration items with version numbers for easy playback and verification.

[0050] Audio observation quality indicators are statistically calculated from audio waveform frame sequences and their timestamp sequences. Audio waveform frame sequences can be acquired by the terminal microphone or read through the system audio interface during call-like sessions. Metadata such as sampling rate, sampling bit width, and number of channels are returned by the system interface and recorded with the session. In specific calculations, the system first performs speech activity detection on the audio stream, dividing the audio into speech and non-speech segments. Speech activity detection can be output by the end-side model or implemented by a general algorithm; its output includes at least the determination of whether each time frame is speech, thus supporting repeatable calculations in subsequent statistics. Based on this, the system calculates short-time energy for each audio frame, which can be obtained by summing the squares of the waveform amplitudes of each frame; the energy statistics of the non-speech segments characterize the noise floor, and the energy statistics of the speech segments characterize the effective speech intensity. The audio signal-to-noise ratio (SNR) is estimated by taking the logarithm of the ratio of the mean energy of the speech segment to the mean energy of the non-speech segment, thus reflecting the changes in signal availability caused by noise interference and end-side gain compression. Simultaneously, the system calculates the proportion of speech segments to all frames within a sliding time window to determine the effective speech segment percentage. This percentage is used to characterize whether there is prolonged silence, strong noise overload, or permission / acquisition anomalies within that time window. Furthermore, it calculates the mean and fluctuation amplitude of the noise floor in non-speech segments to reflect changes in noise patterns caused by environmental noise increases, residual noise reduction at the edge, or automatic gain control. The aforementioned energy, mean, quantiles, and fluctuation amplitude can all be directly obtained from the audio frame energy sequence, with a clear data source and recalculation capability.

[0051] Video observation quality indicators are obtained from the video frame sequence, the timestamp of each frame, and the output statistics of the visual model. The video frame sequence is acquired by the front-facing camera, and the frame rate, resolution, and other configurations are returned by the camera interface. When frame rate fluctuations occur, the system can obtain the actual frame interval and frame rate fluctuation from the difference in timestamps between adjacent frames, serving as an objective basis for the temporal stability of the video. In specific calculations, the system performs face detection and key point localization on each frame, obtaining the face detection results, key point coordinates, and key point confidence. The system determines frames with successful face detection and key point confidence that meet the minimum usability condition as valid frames, and calculates the proportion of valid frames within a sliding time window to characterize the usability of video evidence. The minimum usability condition is not a manually fixed value, but rather an initial lower bound obtained from the quantile statistics of historical stable samples or online stable samples, and is recorded with each version, making the usability / unusability determination under different devices and lighting conditions traceable. Simultaneously, the system performs mean, quantile, and fluctuation statistics on keypoint confidence to characterize the stability of localization and the increased uncertainty caused by occlusion, side profiles, and blur. To reflect the impact of low light and blur on visual evidence, the system can also calculate brightness distribution and gradient statistics for video frames, forming surrogate quantities for sharpness and low light intensity: brightness distribution can be obtained from pixel brightness histograms or mean / quantile statistics, and gradient statistics can be obtained from image edge intensity or gradient magnitude mean statistics. The aforementioned sharpness and low light surrogate quantities are also calculated using a sliding time window to determine their mean and fluctuation, thereby achieving a calculable characterization of degradation states such as low light, blur, and compression artifacts. Face detection results and keypoint confidence are derived from frame-by-frame output of the visual model, while brightness and gradient statistics are directly calculated from image processing, ensuring a clear data source. Text and transcription observation quality indicators are obtained from user-input text, input timestamps, and the return values ​​of the speech transcription service. Speech-to-text transcription can be generated by a local speech recognition module or by a cloud-based speech recognition service. The transcription result should at least include segmented text, corresponding timestamps, and confidence or stability scores to support time window alignment with audio and video evidence and stability statistics. The system reads the confidence or stability scores returned by the transcription service, calculates their mean and fluctuation range within a sliding time window to characterize the reliability of the text evidence; it further calculates anomaly rates such as the proportion of empty transcription segments and the proportion of low-confidence segments to characterize the unusability of text evidence due to weak networks, noise, or recognition failures. All of the above indicators are directly obtained from the transcription service's return values ​​and can be recorded with the conversation for easy verification.

[0052] Link fingerprinting is used to characterize the systematic impact of end-side processing links and transmission transcription links on signal statistical characteristics. This embodiment provides a reproducible data acquisition method and clarifies the source of link statistics and the degradation calculation method when the interface is unavailable. For audio link fingerprinting, the system statistically analyzes changes in short-term energy distribution patterns, frequency band energy distribution ratios, and the degree of noise floor rise in non-speech segments within a sliding time window. This reflects the distortion of prosody and spectral patterns caused by noise reduction, echo cancellation, automatic gain control, and coding compression. The short-term energy distribution pattern can be characterized by changes in energy quantile differences or energy distribution dispersion; the frequency band energy distribution ratio can be obtained by statistically analyzing the energy proportion of each frequency band after multi-band filtering of the audio; and the noise floor rise is directly obtained from the statistics of noise floor in non-speech segments. All the above statistics are calculated from audio waveform frame sequences and do not rely on vendor-specific interfaces. For video link fingerprinting, the system uses frame timestamp differential statistics to analyze frame interval and frame rate fluctuations, reflecting temporal degradation caused by frame drops, stuttering, and weak networks. Simultaneously, the system prioritizes reading encoding parameters and bitrate adaptive change information from media encoding module logs or the system media interface to characterize compression intensity changes. When the terminal cannot provide encoding parameters or log reading is unavailable, the system can use image statistical surrogate quantities to estimate the strength of compression artifacts, such as approximating them through time window statistics of blocky edge anomalies, local high-frequency residual enhancement, or texture distortion indicators. To avoid inconsistencies in link fingerprint caliber due to differences in interface availability across different terminals, this embodiment defines a degradation priority for link statistics: priority is given to reading encoding parameters and bitrate adaptive statistics returned by the system / media framework; when unavailable, frame rate / frame interval fluctuations obtained from video frame timestamp differentials are used as temporal components; based on this, image statistical surrogate quantities such as block effects and high-frequency energy attenuation are used to estimate compression artifact components. The availability markers, degradation paths, and calculation methods used for each component are recorded with the session and written to the version number for playback verification. For transcription link fingerprinting, the system statistically analyzes the fluctuation range of transcription confidence and the proportion of low-confidence segments, and synchronously correlates this with transport layer statistics such as network packet loss rate, latency, and latency jitter to characterize the transcription drift intensity caused by link instability. Network packet loss rate, latency, and jitter can be read and summarized by time windows from the system's network statistics interface, transport layer statistics, and real-time communication statistics, such as the send / receive statistics provided by the real-time session framework. When only partial network statistics are available, the system can still correlate either packet loss rate or jitter with transcription fluctuations to ensure that fingerprint construction does not depend on a single interface. To ensure cross-device comparability, the above-mentioned observation quality indicators and link fingerprints can be normalized using rolling quantile statistics of historical stable samples or online stable samples. Stable samples can be defined as a set of sessions where the overall observation quality indicators meet the usability conditions and the fluctuation is at a low level within a continuous time window; the selection rules for stable samples also record the configuration version number.Normalization boundaries, update frequency, stable sample selection rules, as well as map version number and model version number are written into the configuration or ledger record so that comparison and review can be performed when false alarms, fluctuations or model updates occur.

[0053] Embodiment 3 of the present invention: Based on Embodiment 1, this embodiment further illustrates the implementation method of evaluating the effectiveness of edge-by-edge and adaptively calibrating the edge weights or constraint strength, so that the edge weights or constraint strength can change continuously with the link state to avoid discrete jumps, and ensure that the input and output of edge-by-edge calibration are clear and verifiable.

[0054] In this embodiment, each relation edge in the knowledge graph is recorded in the edge table at least as follows: relation edge identifier, relation type, main dependent modality type, evidence feature binding rule, link fingerprint binding rule, initial edge weight or initial constraint strength, upper and lower limits of edge weight or constraint strength, and a determination rule for whether the edge is active in the current session. The mobile device retrieves the graph version number at the start of the session or when the graph version is updated and caches the edge table fields as needed, recording the version number with the session to ensure the traceability of edge binding rules and edge weight parameters. To avoid meaningless updates to relation edges that are irrelevant to the current session or lack evidence, this embodiment prioritizes performing edge-by-edge evaluation and calibration on a subset of the active relation edges in the current session. The activation determination must at least satisfy the following: the evidence feature bound to the edge can be calculated from the collected modalities in the current session, and the observation quality of the corresponding modality is not lower than the minimum usability condition. If a modality that an edge depends on is missing in the current session, or if the observation quality of that modality is below the minimum availability condition, then that edge will not participate in the calibration update in this session. Only the edge weights or constraint strengths from the previous session will be retained and marked as evidence unavailable or unstable in the session record, in order to avoid long-term contamination of edge weights by degraded data. The minimum availability condition can be determined by the initial lower bound of quantile statistics of historical stable samples or online stable samples, and the version number will be recorded. The data caliber is consistent with that in Example 2.

[0055] During edge-by-edge evaluation, the system first selects the corresponding observation quality score based on the main dependent modality type in the edge table. The observation quality score is obtained according to the method of Example 2. For example, for audio, it can select the comprehensive statistics of audio signal-to-noise ratio estimation and effective speech segment ratio; for video, it can select the effective frame ratio and key point confidence statistics; and for text, it can select the mean of transcription confidence and fluctuation amplitude. If a relation edge depends on two or more modalities, the comprehensive result of multimodal observation quality scores can be selected according to the edge table binding rules, such as taking the lower one or weighting it according to a preset ratio. The comprehensive rules are recorded with the edge table version to ensure that they can be verified. The system selects the corresponding link influence score according to the link fingerprint binding rules in the edge table. The link influence score is also obtained according to the method of Example 2. For example, for audio edges, it can bind statistics such as changes in audio short-term energy distribution morphology, changes in frequency band energy distribution ratio, and noise floor rise; for video edges, it can bind frame rate / frame interval fluctuation and compression artifact proxy; and for text edges, it can bind transcription confidence fluctuation, low confidence segment ratio, and its correlation with network packet loss rate and latency jitter. If the terminal cannot provide partial link statistics interfaces, such as being unable to read encoded parameter logs, then the computable statistical proxy quantity is used as a substitute in the degradation method of Example 2, and the availability status of the current link fingerprint is marked in the session record. To ensure cross-device and cross-session comparability, both the observation quality score and the link impact score can be normalized using rolling quantile statistics of historical stable samples or online stable samples, so that the two are mapped to a unified scale; the normalization boundary, update frequency, and version number are written into the configuration for easy playback and verification. After normalization, a higher observation quality score indicates more reliable evidence, and a higher link impact score indicates a stronger systematic distortion of the statistical form by the link.

[0056] After obtaining the two types of scores, the system calculates the edge-level effectiveness coefficient for each activatable relation edge. To ensure feasibility and recalculability, this embodiment provides an explicit combination method: the edge-level effectiveness coefficient E can be taken as the product of the observation quality score and (1 minus the link influence score), and the result is limited to the range of 0 to 1. This embodiment provides a recalculable approach: let the initial edge weight or initial constraint strength of the edge be W_init, and the upper and lower bounds be W_min and W_max (both recorded according to the map version), then the target strength W_tar can be determined by the following formula: When it is necessary to ensure that prior knowledge is not completely closed, W_min can be set as the lowest preservation lower bound; W_min, W_max, and W_init are sent as edge table fields with the graph version and recorded with the session for easy post-event review. In this way, the effectiveness coefficient continuously decreases when observation quality deteriorates or link influence increases; and continuously recovers when observation quality recovers and link influence weakens. If more conservative suppression is required in actual engineering, nonlinear mapping can be configured in the edge table, such as applying stronger penalties to link influence scores, but the mapping rules should be recorded with each version to avoid creating unverifiable subjective thresholds. After calculating the edge-level effectiveness coefficient, the system performs amplitude limiting and smoothing updates on the edge weight or constraint strength of that edge. In this embodiment, the "target value" of edge weight or constraint strength can be determined by the initial edge weight or initial constraint strength and the edge-level validity coefficient. For example, the initial value can be scaled according to the validity coefficient to obtain the target strength of the edge in this session. When the data is insufficient or it is necessary to ensure that the prior is not completely closed, a minimum lower bound can be set at the same time to ensure that the target strength is not lower than the lower limit configured in the edge table, thereby avoiding the edge weight being weakened to an irrecoverable state in one go.

[0057] The update of edge weights or constraint strengths is implemented smoothly to avoid jumps caused by short-term fluctuations. Smooth updates can be achieved using exponential smoothing: the edge weights or constraint strengths from the previous time step are weighted and fused with the current target strength according to the update rate, so that the new value continuously approaches the target value; furthermore, this embodiment provides the update formula for each relation change: Where α is the update rate, ranging from 0 to 1. To limit a single jump, we can... Implement amplitude limiting, so that ΔW_max represents the maximum single update magnitude and is written to the configuration. W_old represents the edge weight or constraint strength of the relation edge before the current update, and W_new represents the update target value calculated based on the edge-level effectiveness coefficient or long-term adjustment strategy. ΔW represents the update increment. The initial values ​​of α and ΔW_max can be determined from offline historical data with a joint objective of reducing the false positive rate and output volatility. After going live, they can be automatically fine-tuned based on the volatility statistics of stable sessions, and the parameter version number is written to the configuration to ensure the reproducibility of the calibration process. Simultaneously, the maximum change magnitude within a single update time window is set, and a limit is applied to the change amount. The update rate and maximum change magnitude can be initially determined from offline historical data with a joint objective of "reducing the false positive rate and output volatility," and fine-tuned after going live based on the volatility statistics of stable sessions; relevant parameters and version numbers are written to the configuration to ensure the reproducibility of the calibration process. When the observation quality score falls below the minimum availability condition, the system can reduce the update rate of that edge or freeze its update, retaining only the edge weights or constraint strengths from the previous stable state, and marking the reason for freezing in the session record. Smooth updates will resume once the observation quality recovers and the link impact subsides. The minimum availability condition, denoted as Q_min, can be generated from the rolling quantiles of the online stable samples, and its threshold version can be maintained according to the edge's primary dependency mode type. The freezing rule can be recalculated as follows: when... When α=0, make And record the freeze time window and the reason for freezing in the ledger; when Q consecutively meets K time windows Furthermore, when the link impact score L falls back to the normal range, α is restored to the normal range and the update continues smoothly according to the above update formula. Q is the lowest usability judgment score / availability score for the current time window, with a value of 0 to 1, calculated from the observation quality and link impact of the current window. For example, the usability score is... ,in (where L is the link impact score, representing the observed quality score). Q_min: Freeze threshold, obtained from the rolling quantiles of the stable samples. Q_rec: Thaw recovery threshold (hysteresis upper edge), used for jitter reduction, satisfying... , can be set to (Δ_Q represents the recovery margin, written to the configuration, and recorded in the version). Both K and Q_rec are written to the configuration and their version numbers are recorded to avoid edge weight jitter caused by frequent freezing / unfreezing. This freezing strategy, like the minimum availability condition, is generated from historical statistics or online quantile statistics and recorded with each version.

[0058] Through the aforementioned edge-by-edge evaluation and smoothing calibration mechanism, the system can perform measurable and traceable continuous calibration of the validity of knowledge graph relationship edges under conditions such as changes in end-side processing links, weak network transmission, and transcription fluctuations, providing stable input for subsequent edge-level conflict calculation, conflict attribution, and feedback adjustment.

[0059] Embodiment 4 of the present invention: Based on Embodiment 1, this embodiment further explains the calculation object, normalization method and attribution discrimination criterion of the edge-level conflict quantity, so that the system can distinguish between observation degradation and relation failure and provide verifiable evidence for the source of conflict during post-event playback.

[0060] In this embodiment, the system generates edge evidence quantity and edge inference quantity for at least some relationship edges, and calculates edge-level conflict quantity within a sliding time window. The sliding time window length and update frequency follow the configuration method of Embodiment 2, with initial values ​​determined by the fluctuation statistics of historical stable samples or online stable samples, and the version number is recorded in the configuration for easy playback and review. The edge evidence quantity is obtained by mapping the evidence features bound to the edge. The evidence features are derived from the modal feature extraction results of Embodiment 1 S1, specifically including but not limited to: speech rate, pause ratio, fundamental frequency and energy fluctuation correlation statistics; facial expression intensity, facial movement amplitude, gaze stability; negative word ratio, negative expression ratio, emotional tendency intensity, etc. Each relationship edge records its evidence feature binding rules in the edge table, clarifying which features should constitute the edge evidence quantity and the combination caliber of the features, thereby ensuring that the same relationship edge is counted consistently in different sessions. In this embodiment, the edge evidence quantity is uniformly mapped to an evidence strength scale of 0 to 1, where 0 indicates that the evidence for the symptom / tendency described by the edge is weak in the current session, and 1 indicates strong evidence. The mapping method can adopt one of the following recalculated approaches: determine the low and high quantile boundaries of the evidentiary feature using quantile statistics of historical labeled samples or online stable samples, and truncate and linearly normalize the current feature value; or, in a follow-up scenario, first establish an individual baseline interval, for example, statistically determine the stable range of the feature in multiple consecutive high-quality conversations, and then map the degree of deviation from the individual baseline as the evidence strength. The generation rules, update frequency, and version number of the above quantile boundaries or individual baselines are all written into the configuration or ledger to ensure verifiability. For example, speech rate and pause ratio can be combined according to the side table binding rules to form the lag evidence strength; facial expression intensity and facial movement amplitude can be combined to form the low activity evidence strength; the proportion of negative words and the proportion of negative expressions can be combined to form the negative cognitive evidence strength. The mapping rules are stored with the graph side table version, and consistent results can be obtained by recalculating according to the same version of the rules during playback.

[0061] The edge inference quantity is used to express the tendency of the graph prior to constrain the edge evidence in the current session. Its data sources include: the relationship type and direction of the edge, such as positive association, negative association, conditional association, etc., the edge weight or constraint strength after calibration in Example 3, and the current state estimate or neighborhood evidence summary value of the node associated with the edge in this session. To ensure feasibility, this embodiment provides a recalculated inference generation method: the system first summarizes the upstream nodes or neighborhood nodes related to the edge according to the graph structure to obtain the prior support strength for the target node; then, it performs a positive or negative mapping of the support strength in combination with the relationship direction of the edge, and scales the mapping result according to the edge weight or constraint strength obtained in Example 3, finally obtaining an inference strength scale of 0 to 1. The upstream / neighborhood aggregation can be achieved by weighted averaging or weighted voting of neighborhood evidence, with the weighting coefficients coming from graph edge weights or node confidence. When neighborhood evidence is missing or of low quality, the edge inference can degenerate to the initial constraint tendency of that edge, for example, falling back to a neutral level, and the reason for the degeneration is recorded in the ledger, thereby avoiding instability caused by unavailable data.

[0062] The edge-level conflict quantity is used to measure the degree of inconsistency between the edge evidence quantity and the edge inference quantity. To suppress spurious conflicts caused by noise and degradation, this embodiment introduces the observation quality score corresponding to the edge's main dependent mode when normalizing the conflict quantity. The observation quality score is taken from the caliber of Embodiment 2, and can be the lower value or a proportional synthesis in multimodal dependence according to the edge table rules.

[0063] This embodiment provides a clear definition: the edge-level conflict quantity can be obtained by dividing the absolute value of the difference between the edge evidence quantity and the edge inference quantity by the sum of the observation quality score and the minimum positive constant, i.e., edge-level conflict quantity = |edge evidence quantity − edge inference quantity| ÷ (observation quality score + minimum positive constant). The minimum positive constant is used to avoid numerical instability when the observation quality score approaches zero; it can be a low quantile value of the stable sample observation quality score or a preset small constant, and is recorded with each version. Through this normalization method, when the observation quality is low, the denominator becomes smaller, potentially amplifying the conflict quantity. Therefore, this embodiment further includes a lower bound gating: when the observation quality score is below the minimum usable condition, the conflict quantity is not used for relation failure judgment; it is only allowed to enter the observation degradation branch, thus ensuring that low-quality sessions do not mistakenly trigger relation failure handling. Specifically, when… When this occurs, the system marks the edge as a candidate for degradation. Its conflict amount is only used to suppress the modal contribution coefficient of short-term feedback regulation, is not included in the cumulative failure amount, and does not trigger a weakening update of long-term feedback regulation; only when Furthermore, an edge is only allowed to enter the relationship failure judgment and long-term feedback adjustment handling branch when the conflict quantity meets the persistence condition. Q_min and Q_norm are maintained by rolling quantile statistics and their version numbers are recorded, thus avoiding fixed manual settings that lead to unverifiable results. Regarding conflict source attribution, the system jointly judges edge-level conflict quantity and observation quality changes, introducing two types of conditions: synchronicity and persistence. When the edge-level conflict quantity significantly increases within a certain time window, and the observation quality score of the mode that the edge mainly relies on synchronously decreases and falls below the minimum usability condition within the same time window, the system attributes the conflict to observation degradation. Synchronicity can be determined by simultaneously satisfying both conflict increase and quality decrease within the same sliding time window, and the start and end times, conflict quantity, and quality score of that time window are recorded in the ledger to ensure replay verifiability. When the observation quality score remains within the normal range for a long period, not below the minimum usability condition, and above the normal range threshold, while the edge-level conflict quantity remains high or its cumulative value exceeds the persistence threshold within multiple consecutive time windows, the system attributes the conflict to relationship failure. Persistence can be assessed using one of the following recalculated criteria: the number of conflicts exceeds the high conflict threshold for several consecutive time windows; or the proportion of high conflict time windows exceeds the proportion threshold within a fixed-length cumulative window; or a weighted cumulative calculation of conflict amounts is performed, and failure is determined when the cumulative value exceeds the threshold and the observation quality remains normal. The minimum availability condition, normal interval threshold, high conflict threshold, proportion threshold, or cumulative threshold can all be generated from historical statistics or online rolling quantile statistics and updated periodically. The threshold generation rules, window length, update frequency, and version number are written into the configuration.

[0064] Embodiment 5 of the present invention: Based on Embodiment 1, this embodiment further explains the triggering conditions, update methods and cooling hysteresis mechanisms of short-term feedback regulation and long-term feedback regulation, and further explains the quality-weighted smoothing caliber of the follow-up trend, so as to reduce the interference of short-term degradation on the judgment of long-term trends.

[0065] When conflicts are attributed to observation degradation, the system triggers short-term feedback adjustment, prioritizing a reduction in the contribution of degraded modes to the fusion output within a single session or short time window. The contribution update is determined jointly by the magnitude of the observation quality degradation of the mode and the amount of conflict on its associated edges, while a lower bound is set and a smooth update is employed. The lower bound of the contribution can be obtained statistically from historical stable samples. Let the contribution of mode m be... ∈ , This is the lower bound (0-1) for the minimum contribution of mode m, used to ensure that the mode is not completely shut down when available. It is issued by the configuration table according to mode type and the version number is recorded. This can be obtained from statistical analysis of stable online samples, for example, by taking the low quantile of the contribution distribution of the mode during the stable period, or by setting it as a constant according to business security policies. The associated conflict summary is... The quantile or mean of the modal correlation edge conflict quantity within the time window can be taken, and the target contribution can be set as follows: And adopt smooth updates ; where κ and β are configuration parameters and record version numbers, which represent the contribution (state quantity) of the previous window. The first window is obtained using explicit initialization rules, and each subsequent window is derived from the previous window. Direct inheritance.

[0066] When a conflict is attributed to a relationship failure, the system triggers long-term feedback adjustment. Only under the condition that "observation quality remains normal and conflicts consistently meet the threshold" will the edge weights or constraint strengths of the corresponding relationship edges be weakened in small steps. If the observation quality is unstable or below the usable lower bound, long-term feedback adjustment updates are paused, and only short-term feedback adjustment is retained to temporarily suppress the modal contribution coefficient. Long-term feedback adjustment sets a cooling-off time window: within the cooling-off window, the weakening of the same relationship edge is not repeatedly triggered; hysteresis recovery is also set: when subsequent observation quality recovers and the conflict amount is below the recovery threshold within a continuous time window, the edge weights or constraint strengths are allowed to smoothly return to a reasonable range. The cooling-off time window, persistence threshold, and recovery threshold can all be generated from historical statistics or online rolling quantile statistics and written to the configuration record version for verification. Regarding the generation of follow-up trends, the trend is formed from a risk score sequence of consecutive sessions, and quality-weighted smoothing is performed based on the observation quality indicators of each session. When a session suffers from poor observation quality due to weak network conditions, low light, occlusion, frame drops, or unstable transcription, its weight is reduced, thus suppressing its contribution to the trend. When session quality recovers and the number of conflicts decreases, the weight increases, the trend output smoothly reverts, and it more accurately reflects the direction of risk changes. Smoothing can be achieved using a sliding time window weighted average or exponential smoothing. The window length or smoothing coefficient can be initially selected from targets in historical data with small trend fluctuations and no sluggishness to real changes. After going live, it can be fine-tuned based on stable sample fluctuation statistics, and the version can be recorded.

[0067] Embodiment Six of the present invention: Based on Embodiment One, this embodiment further explains the fields, writing timing, sorting and positioning rules, and hysteresis and cooling conditions for rollback recovery of the edge-level ledger, so that the system can locate a few key failure relationships and support smooth recovery, reduce maintenance costs and improve traceability.

[0068] In this embodiment, the system establishes an edge-level ledger for at least some relationship edges. The ledger records at least: relationship edge identifier, relationship type, main dependent mode type, current link fingerprint summary, corresponding observation quality index summary value, edge-level conflict quantity, conflict attribution result, edge weight / constraint strength values ​​before and after calibration, triggered short-term / long-term feedback adjustment actions and trigger time windows, threshold version number, and graph version number. Ledger entries are written when: a relationship edge undergoes calibration update, a long-term feedback adjustment weakening update, a hysteresis recovery, or a trigger rollback action, to ensure key changes are traceable. The system performs a cumulative evaluation of edge-level conflict quantity under cross-session conditions and outputs a set of key conflicting relationship edges for review and location. The cumulative evaluation can use one or a combination of the following criteria: cumulative number of times exceeding the threshold, high conflict percentage within the sliding window, and joint screening of persistent conflicts under normal observation quality conditions. Relevant thresholds and windows can be generated from historical statistics or online rolling quantile statistics and written into the configuration record version. Regarding rollback and recovery, when a relation edge experiences persistently high conflict within the normal observation quality range, and long-term feedback adjustments fail to suppress abnormal output, the system can trigger a rollback operation. This returns the edge weight or constraint strength to the previous stable range or the previous stable graph version, and records the rollback operation and its corresponding version number in the log. Simultaneously, the system employs a recovery mechanism: when the observation quality index meets the recovery conditions, and the edge-level conflict amount remains below the recovery threshold within a continuous time window, the edge weight or constraint strength is allowed to smoothly return to a more reasonable range. To avoid edge weight oscillations, a cooling-off period is set after rollback or recovery, during which similar actions are not repeatedly triggered. The recovery conditions, recovery threshold, and cooling-off period length can all be generated and recorded from historical statistics or online rolling quantile statistics for review and playback.

Claims

1. A multimodal feature fusion and recognition method integrating knowledge graph constraints, characterized in that, include: S1: Acquire multimodal data generated by the target object during natural interaction, perform feature extraction on each modal data to obtain modal features for mental health status risk assessment, and calculate observation quality indicators based on the multimodal data and the objective information of its acquisition and transmission process. The multimodal data includes at least two or more of the following: audio data, video data, and text data. The observation quality index is used to characterize the reliability of the corresponding modal observations; S2: Based on the statistical morphology and link statistics of the multimodal data, calculate the link fingerprint and obtain a knowledge graph related to mental health status identification; The link fingerprint is used to characterize the intensity of the systematic influence of the end-side processing link and the transmission transcription link on the signal statistical characteristics; The knowledge graph contains relation edges associated with the modality features; S3: Based on the link fingerprint and the observation quality index, evaluate the effectiveness of one or more of the relation edges one by one and adaptively calibrate the edge weight or constraint strength of the corresponding relation edges; S4: Perform multimodal feature fusion reasoning based on the modal features and the calibrated knowledge graph relationship edges; The fusion inference sets a modal contribution coefficient for each modal feature, which is used to characterize the weighted contribution ratio or gating ratio of each modal feature to the fusion output; In the fusion reasoning process, for one or more of the relation edges, edge-level conflict quantity is calculated edge by edge to quantify the degree of inconsistency between the graph constraint inference and the multimodal evidence, and the conflict source is attributed based on the edge-level conflict quantity and the observation quality index to distinguish between conflicts caused by observation degradation and conflicts caused by relation failure. S5: Perform feedback adjustment based on the attribution results, wherein conflicts caused by observation degradation correspond to the adjustment of the modal contribution coefficient, and conflicts caused by relation failure correspond to the adjustment of relation edge weights or constraint strength. S6: Based on the fusion inference results after calibration and feedback adjustment, output the mental health status recognition result; The identification results include at least a risk score, risk level, or follow-up trend; The follow-up trend refers to the follow-up trend of the risk score and / or the follow-up trend of the risk level.

2. The method as described in claim 1, characterized in that, The observation quality indicators include audio observation quality indicators, video observation quality indicators, text observation quality indicators, and transmission observation quality indicators.

3. The method as described in claim 1, characterized in that, The link fingerprint consists of at least one signal statistical morphology index and at least one link statistical index; The signal statistical morphology indicators include one or more of the following: changes in the quantile difference of audio short-time energy distribution, changes in audio frequency band energy ratio, and statistics on the noise floor of non-speech segments. The link statistics include one or more of the following: coding parameter change statistics, frame rate fluctuation statistics, bit rate adaptive statistics, and speech transcription confidence statistics.

4. The method as described in claim 1, characterized in that, The process of evaluating the effectiveness of each edge and adaptively calibrating the edge weights or constraint strengths of the corresponding relational edges includes: Based on the link fingerprint and observation quality index, the edge-level effectiveness coefficient is calculated for each relation edge, and the corresponding edge weight or constraint strength is updated by limiting or smoothing according to the edge-level effectiveness coefficient, so that the edge weight or constraint strength is continuously adjusted with the link state change rather than discretely jumping.

5. The method as described in claim 1, characterized in that, The edge-level conflict quantity is obtained by normalizing the difference between the graph constraint inference results and the observations corresponding to the multimodal evidence; The graph constraint inference result is as follows: In S4, based on the relation edge weights or constraint strengths calibrated in S3, the inference quantity obtained by performing inference or calculating the graph constraint terms on the knowledge graph is obtained. The observations corresponding to the multimodal evidence are: in S4, the amount of evidence obtained by mapping or aggregating the modal features extracted in S1 according to the evidence binding rules of the corresponding relation edges; The normalization is at least associated with the observation quality index; The association includes: for at least one modality bound to the relation edge, determining a normalized scale parameter or weight coefficient based on its observation quality index, and using the scale parameter to scale or weight the difference so that the contribution of the normalized difference is suppressed when the observation quality decreases.

6. The method as described in claim 1, characterized in that, The conflict source attribution includes: calculating degradation indicators and failure indicators based on the edge-level conflict quantity and observation quality index, and determining the conflict source based on the relative magnitude of the degradation indicators and failure indicators; The degradation indicator is used to characterize the increased inconsistency caused by a decrease in observation reliability; The failure indicator is used to characterize the degree to which inconsistencies persist under the condition that the reliability of observations remains normal.

7. The method as described in claim 1, characterized in that, The feedback regulation includes: short-term feedback regulation and long-term feedback regulation; The short-term feedback adjustment rapidly updates the modal contribution coefficient based on the observation quality index within a single session or short time window to suppress output fluctuations caused by weak networks, low light, or occlusion. The long-term feedback adjustment cumulatively evaluates the amount of edge-level conflict under cross-session conditions and adjusts the edge weights or constraint strengths of the relation edges when the persistence condition is met.

8. The method as described in claim 1, characterized in that, The method also includes: establishing a side-level ledger to achieve closed-loop traceable adjustment; The edge-level ledger records at least the relation edge identifier, link fingerprint, observation quality index, edge-level conflict quantity, and corresponding calibration and feedback adjustment actions. Based on the cumulative results of the edge-level conflict quantity, the relation edges are sorted and located to output a set of key conflict relation edges for review and operation and maintenance location.

9. The method as described in claim 8, characterized in that, The edge-level ledger is further used to trigger rollback or recovery; The rollback or recovery satisfies the hysteresis and cooling conditions; The hysteresis and cooling conditions include at least the observation quality index meeting the recovery conditions and the edge-level conflict amount being below the recovery threshold within a continuous time window. The recovery conditions and the recovery threshold are generated by historical statistics or online quantile statistics.

10. The method as described in claim 1, characterized in that, The mental health status identification result is the depression risk identification result; The identification results include depression risk score, depression risk level or follow-up trend, and the knowledge graph is a depression-related knowledge graph or a subgraph selected from a depression-related knowledge graph. The follow-up trend refers to the follow-up trend of the risk score and / or the follow-up trend of the risk level.

Citation Information

Patent Citations

  • Depression detection system based on social media

    CN113139062A

  • Depression recurrence risk intervention method, device, equipment and medium

    CN120413024A