Collection method and system for multi-mode depressive emotion data
By employing hardware triggering mechanisms and time deviation compensation technology, the synchronization problem in the collection of multimodal depressive mood data was solved, achieving data consistency and high-quality collection over time.
Patent Information
- Application Number
- CN202511964105.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies cannot accurately and synchronously collect multimodal depressive mood data, which increases the difficulty of subsequent analysis.
A hardware triggering mechanism is used to synchronously start video and audio data acquisition. A unified time reference is constructed using a high-precision monotonic clock to calculate and compensate for time deviations, thereby achieving synchronization of video and audio data.
Ensure consistency of data from different modalities over time, suppress nonlinear drift, and improve data quality.
Smart Images

Figure CN121622041A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data collection, and more particularly to a multi-modal depression emotion data collection method and system. BACKGROUND
[0002] At present, the research of affective computing and mental health relies on the collection and analysis of multi-modal data, especially when judging the intensity of depression emotion, facial expression, speech signal and gait information as the key manifestation of depression emotion, the efficient and accurate synchronous collection is particularly important.
[0003] The prior art proposes a multi-modal time sequence processing depression state data processing method, electronic equipment and medium, first, the facial information, head posture information and eye data of the subject are collected, and based on the time sequence configuration, the multi-modal data is configured; then the multi-modal data is preprocessed, and the multi-modal data set is made and loaded into the trained depression data recognition neural network model to obtain the depression data processing result. The scheme based on the deep learning algorithm processes the collected multi-modal depression emotion data, provides an automatic data processing and analysis method, provides data reference for psychiatrists, and improves the processing efficiency, but the collected data cannot accurately start, run and record multiple sensors at the same time point. Even if all the devices start at the same time, due to the slight difference of their own clock or system processing jitter, their time axis will gradually lose synchronization in the long recording process, which greatly increases the difficulty of subsequent data analysis. SUMMARY
[0004] In order to solve the problem that the existing multi-modal depression emotion data collection method cannot synchronously collect the data of different modalities of the depression emotion subject to be analyzed, the present application proposes a multi-modal depression emotion data collection method and system, which collects multi-modal depression emotion data and synchronously processes it, ensures the synchronization of data between different modalities, and improves the quality of the collected data.
[0005] In order to achieve the above technical effects, the technical scheme of the present application is as follows: In a first aspect, the present application proposes a multi-modal depression emotion data collection method, comprising the following steps: Synchronously starting the collection of video data and audio data based on a hardware trigger mechanism; the video data includes facial video data of the subject and gait video data of the subject, and the audio data includes questionnaire recording audio data of the subject and reading audio data of the subject; A unified time reference is constructed, the time deviation of the video data and the audio data is calculated based on the preset sampling parameters of the video data and the audio data, the unified time reference, and the collected video data and audio data are drift compensated based on the time deviation. The video data and the audio data after the drift compensation are taken as the final multi-modal depression emotion data, and the collection of the multi-modal depression emotion data is completed.
[0006] In the technical solution, first, the collection of the video data and the audio data is started synchronously based on a hardware trigger mechanism, and the time axes of the video data and the audio data are aligned; a high-precision monotonic clock is taken as a unified time reference, the time deviation of the video data and the audio data is calculated based on the unified time reference, a preset video data target frame rate and a preset audio data sampling rate, the collected video data and audio data are drift-compensated based on the time deviation, and finally the video data and the audio data after the drift compensation are taken as the final multi-modal depression emotion data, and the collection of the multi-modal depression emotion data is completed. The hardware synchronization mechanism is adopted to ensure that all the collection units start working at a unified time, guarantee the consistency of the data between different collection units on the time axis, and effectively inhibit the nonlinear drift of the video data and the audio data in a long-time collection by using the drift compensation, so that the video data and the audio data are synchronized, and the quality of the collected data is improved.
[0007] Preferably, when the collection of the video data and the audio data is started synchronously based on the hardware trigger mechanism, all the collection units are started synchronously based on a unified hardware clock, and the collection of the video data and the audio data is started; the collection units include a forward video collection unit, a lateral video collection unit and an audio collection unit. The facial video data of the subject is collected by the forward video collection unit, the gait video data of the subject is collected by the forward video collection unit and the lateral video collection unit cooperatively, and the audio data is collected by the audio collection unit.
[0008] Preferably, the preset sampling parameters include a video target frame rate and an audio sampling rate ; the process of calculating the time difference of the video data and the audio data is as follows: the video time is calculated based on a starting time stamp of the starting of the forward video collection unit, a video frame sequence number and the video target frame rate , and the expression is as follows:
[0009] the audio time is calculated based on a starting time stamp of the starting of the audio collection unit, an audio sampling point number and the audio sampling rate , and the expression is as follows:
[0010] Based on video time and audio time Calculate the time deviation between audio and video times. The expression is: .
[0011] Preferably, based on the time deviation, drift compensation is performed on the acquired video and audio data, the process of which is as follows: Determine the time deviation between audio and video times If the time difference is less than the preset first time difference threshold, then video and audio data synchronization will not be performed. If the time difference between audio time and video time If the sampling rate is greater than or equal to a preset first time difference threshold and less than a preset second time difference threshold, then adjust the sampling rate. Until time deviation Less than the first time difference threshold; If the time difference between audio time and video time If the time difference is greater than or equal to the second time difference threshold, synchronization is performed based on the deviation direction between the video and audio data. If the video data lags behind the audio data, non-critical frames in the video data are discarded until the time difference is reached. If the time difference is less than the first time difference threshold; if the video data leads the audio data, then perform frame interpolation on the non-key frames in the video data until the time difference is reached. Less than the first time difference threshold.
[0012] Preferably, the video data acquisition process further includes: The candidate regions for the subject's face and gait are calculated, and facial and gait video data of the subject are collected based on these candidate regions. The process is as follows: The two-dimensional coordinates of the head in the lateral image of the subject are extracted in real time using the lateral video acquisition unit. Combined with the pre-calibrated camera parameters, the two-dimensional coordinates of the head are used to calculate the estimated position of the head in three-dimensional space. Based on the pre-calibrated spatial relationship matrix between the forward video acquisition unit and the lateral video acquisition unit, the three-dimensional spatial position of the head is projected onto the imaging plane of the forward video acquisition unit, and the candidate regions in the imaging plane of the forward video acquisition unit where the subject's face appears are calculated. Within the candidate region, a first facial capture model preset in the forward video capture unit is invoked to capture facial expressions, and the confidence level of the facial expression capture result is obtained. The confidence level is compared with a preset confidence threshold. If the confidence level is higher than the threshold, the facial expression capture result is output as video data; if the confidence level is lower than the threshold, a second facial capture model is switched within the candidate region for detection, and the facial expression capture result is output as video data.
[0013] Preferably, the video data acquisition process further includes: evaluating the quality of the acquired video data, the process of which is as follows: Based on the Euler angles of the subject's head, the subject's posture compliance index was calculated, and based on the brightness distribution of the subject's facial area, the illumination compliance index was calculated. The posture compliance index and the illumination compliance index are compared with preset quality thresholds. If both indexes are greater than the preset quality thresholds, video data acquisition is continued. If either index is lower than the preset quality threshold, the subject is guided to adjust their posture until both indexes are greater than the preset quality thresholds, and then video data acquisition is resumed. Based on the statistical analysis and guidance records of indicators during the video data acquisition process, data quality level labels are generated for the video data. The data quality level labels are then associated with and stored with the video data to complete the acquisition of multimodal depressive mood data.
[0014] Preferably, the audio data acquisition process further includes: using a callback function to receive real-time audio data, and using a mutex lock to prevent other operations from writing the audio data during the receiving process.
[0015] Secondly, this application also proposes a data acquisition system for multimodal depressive mood data, the system comprising: The synchronous data acquisition module is used to synchronously start the acquisition of video data and audio data based on a hardware triggering mechanism; the video data includes the subject's facial video data and the subject's gait video data, and the audio data includes the subject's questionnaire recording audio data and the subject's reading audio data. The drift compensation module is used to construct a unified time reference. Based on the preset sampling parameters of video data and audio data and the unified time reference, it calculates the time deviation of video data and audio data, and performs drift compensation on the acquired video data and audio data based on the time deviation. The data storage module is used to collect multimodal depressive mood data by using the drift-compensated video and audio data as the final multimodal depressive mood data.
[0016] Thirdly, this application also proposes a computer device, which includes a memory, a processor, and a computer program stored in the memory that can be run on the processor. The processor executes the computer program to implement the aforementioned method for collecting multimodal depressive mood data.
[0017] Fourthly, this application also proposes a computer storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to execute the aforementioned method for collecting multimodal depressive mood data.
[0018] Compared with the prior art, the beneficial effects of the present invention are: This invention proposes a method and system for acquiring multimodal depressive mood data. First, it synchronously initiates the acquisition of video and audio data based on a hardware trigger mechanism, aligning the timelines of the video and audio data. Using a high-precision monotonic clock as a unified time reference, and based on this unified time reference, a preset target frame rate for video data, and a preset sampling rate for audio data, the time deviation between the video and audio data is calculated. Based on this time deviation, drift compensation is applied to the acquired video and audio data. Finally, the drift-compensated video and audio data are stored according to data type, completing the acquisition of multimodal depressive mood data. This application employs a hardware synchronization mechanism to ensure that all acquisition units start working at the same time, guaranteeing the consistency of data on the timeline between different acquisition units. Furthermore, by utilizing drift compensation, it effectively suppresses nonlinear drift of video and audio data during long-term acquisition, achieving synchronization of video and audio data and improving the quality of the acquired data. Attached Figure Description
[0019] Figure 1 A flowchart illustrating the method for collecting multimodal depressive mood data proposed in Embodiment 1 of the present invention; Figure 2 This diagram illustrates the layout of the data collection facilities for collecting multimodal depressive mood data, as proposed in Embodiment 2 of the present invention. Figure 3 This is a schematic diagram of the structure of a data acquisition system for multimodal depressive mood proposed in Embodiment 3 of the present invention; Figure 4 This is a schematic diagram of the structure of the computer device proposed in Embodiment 4 of the present invention. Detailed Implementation
[0020] The accompanying drawings are for illustrative purposes only and should not be construed as limiting the scope of this patent. To better illustrate this embodiment, some parts of the accompanying drawings may be omitted, enlarged, or reduced, and do not represent the actual dimensions; It is understandable to those skilled in the art that some well-known details may be omitted from the accompanying drawings.
[0021] The technical solution of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0022] The positional relationships depicted in the accompanying drawings are for illustrative purposes only and should not be construed as limiting this patent. Example 1 This embodiment proposes a method for collecting multimodal depressive mood data. A flowchart illustrating this method can be found here. Figure 1 This includes the following steps: S1. The acquisition of video and audio data is initiated simultaneously based on a hardware triggering mechanism; the video data includes the subject's facial video data and the subject's gait video data, and the audio data includes the subject's questionnaire recording audio data and the subject's reading audio data. S2. Construct a unified time reference. Based on the preset sampling parameters of video data and audio data and the unified time reference, calculate the time deviation of video data and audio data. Based on the time deviation, perform drift compensation on the collected video data and audio data. S3. Use the drift-compensated video and audio data as the final multimodal depressive mood data to complete the collection of multimodal depressive mood data.
[0023] In this embodiment, video and audio data acquisition is first initiated synchronously based on a hardware triggering mechanism, aligning the timelines of the video and audio data. A high-precision monotonic clock is used as a unified time reference. Based on this unified time reference, the target frame rate and frame number of the video data, and the sampling rate and number of sampling points of the audio data, the video and audio timelines are determined, and their time deviation is calculated. Drift compensation is then performed on the acquired audio and video data based on this time deviation to achieve cross-modal temporal alignment. Finally, the drift-compensated video and audio data are stored according to data type, completing the acquisition of multimodal depressive mood data. This application employs a hardware synchronization mechanism to ensure that all acquisition units start working at the same time, guaranteeing the consistency of data on the timeline between different acquisition units. Furthermore, by utilizing drift compensation, nonlinear drift of video and audio data is effectively suppressed during long-term acquisition, achieving synchronization of video and audio data and improving the quality of the acquired data.
[0024] Example 2 In this embodiment, when the acquisition of video and audio data is simultaneously initiated based on the hardware triggering mechanism, all acquisition units are simultaneously started based on a unified hardware clock to initiate the acquisition of video and audio data; the acquisition units include a forward video acquisition unit, a side video acquisition unit, and an audio acquisition unit; The subject's facial video data was acquired through a forward video acquisition unit, the subject's gait video data was acquired through a combination of a forward video acquisition unit and a lateral video acquisition unit, and the audio data was acquired through an audio acquisition unit.
[0025] Specifically, each of the forward video acquisition unit, the side video acquisition unit, and the audio acquisition unit has a built-in independent timestamp marking module. This module starts the acquisition operation of all acquisition units through a hardware synchronization signal, ensuring that all devices start working at the same time. During the acquisition process, all acquisition units continuously record and mark the timestamp of each frame and align the data with the unified timestamp of the system to ensure accurate data synchronization between different devices.
[0026] Specifically, the layout diagram of the acquisition unit is as follows: Figure 2 As shown, Figure 2 In this setup, both the forward video acquisition unit 1 and the lateral video acquisition unit 3 use cameras as hardware, while the audio acquisition unit 2 uses a microphone. In the space where data on the subject's depressive mood is collected, the forward video acquisition unit 1 is positioned in front of the computer screen, directly in front of the subject, approximately 0.8 to 1 meter away. The audio acquisition unit 2 is placed next to the forward video acquisition unit 1, approximately 20 to 30 centimeters away from the subject, to capture the subject's speech signal with high fidelity. The lateral video acquisition unit 3 is mounted on one wall of the room, approximately 1.7 meters above the ground. The acquisition angle of the lateral video acquisition unit 3 is parallel to the subject's gait trajectory, approximately 2 to 2.5 meters from the walking baseline. This layout ensures that the forward video acquisition unit 1 can effectively acquire the subject's facial expression features, while the lateral video acquisition unit 3 can completely record the subject's gait information and provide the lateral perspective data required for skeletal estimation. Through this coordinated forward and lateral arrangement, the system can achieve simultaneous multimodal data acquisition and face locking under special behaviors such as subjects looking down, tilting their heads, or avoiding the camera.
[0027] Specifically, the facial video data reflects changes in the subject's facial expression features. More specifically, after the subject confirms the start of the data collection, an emotional stimulation video is played on the display screen, and the forward video acquisition unit 1 is activated to collect the subject's facial expression features. The questionnaire audio data reflects the subject's tone of voice when answering the questionnaire questions. After the facial video data collection ends, the display screen pops up questionnaire recording dialog boxes in sequence according to gender and age group, and the audio data acquisition unit 2 collects the subject's voice when answering the questionnaire questions. Similarly, the reading audio data reflects the subject's tone of voice when reading aloud. After the questionnaire audio data collection ends, the display screen pops up a text fragment, and the audio data acquisition unit 2 collects the subject's voice when reading the text fragment. The gait video data reflects the subject's limb movement features when walking. After the reading audio data collection ends, the display screen guides the subject to walk to collect gait video data, which is collected based on the forward video acquisition unit 1 and the lateral video acquisition unit 3. The video and audio data are collected sequentially in the order of facial video data, questionnaire audio recordings, reading aloud audio data, and gait video data. When an abnormality occurs at a certain stage, the system automatically recovers to the most recent controllable stage, reducing the error rate and improving process traceability.
[0028] In this embodiment, the preset sampling parameters include the target video frame rate. and audio sampling rate The process of calculating the time difference between video and audio data is as follows: Based on the start timestamp when the forward video acquisition unit starts Video frame sequence number and video target frame rate Calculate video time The expression is:
[0029] Based on the start timestamp when the audio acquisition unit starts Number of audio sampling points and audio sampling rate Calculate audio time The expression is:
[0030] Based on video time and audio time Calculate the time deviation between audio and video times. The expression is: .
[0031] Specifically, the preset sampling parameters also include the bit depth of the audio acquisition unit, wherein the audio acquisition unit uses a PCM16 bit depth and The sampling rate is used to collect audio data; the target frame rate of the video... .
[0032] In this embodiment, drift compensation is performed on the acquired video and audio data based on the time deviation. The process is as follows: Determine the time deviation between audio and video times If the time difference is less than the preset first time difference threshold, then video and audio data synchronization will not be performed. If the time difference between audio time and video time If the sampling rate is greater than or equal to a preset first time difference threshold and less than a preset second time difference threshold, then adjust the sampling rate. Until time deviation Less than the first time difference threshold; If the time difference between audio time and video time If the time difference is greater than or equal to the second time difference threshold, synchronization is performed based on the deviation direction between the video and audio data. If the video data lags behind the audio data, non-critical frames in the video data are discarded until the time difference is reached. If the time difference is less than the first time difference threshold; if the video data leads the audio data, then perform frame interpolation on the non-key frames in the video data until the time difference is reached. Less than the first time difference threshold.
[0033] Specifically, the range of the first time difference threshold is 40ms to 80ms. Preferably, the first time difference threshold is 50ms, because the human ear's perception threshold for audio-visual asynchrony is usually around 100ms. Setting the first time difference threshold to 50ms can ensure that neither the user nor the algorithm can perceive the slight jitter when compensation is not triggered, thus guaranteeing the basic acquisition quality.
[0034] Specifically, the second time difference threshold ranges from 150ms to 300ms, preferably 200ms. This is because when the time difference exceeds 200ms, the alignment of facial expressions and audio features will deviate significantly, affecting the accuracy of the depression recognition model. In this case, forced corrections such as frame dropping or frame interpolation must be implemented.
[0035] In this embodiment, the video data acquisition process further includes: The candidate regions for the subject's face and gait are calculated, and facial and gait video data of the subject are collected based on these candidate regions. The process is as follows: The two-dimensional coordinates of the head in the lateral image of the subject are extracted in real time using the lateral video acquisition unit. Combined with the pre-calibrated camera parameters, the two-dimensional coordinates of the head are used to calculate the estimated position of the head in three-dimensional space. Based on the pre-calibrated spatial relationship matrix between the forward video acquisition unit and the lateral video acquisition unit, the three-dimensional spatial position of the head is projected onto the imaging plane of the forward video acquisition unit, and the candidate regions in the imaging plane of the forward video acquisition unit where the subject's face appears are calculated. Within the candidate region, a first facial capture model preset in the forward video capture unit is invoked to capture facial expressions, and the confidence level of the facial expression capture result is obtained. The confidence level is compared with a preset confidence threshold. If the confidence level is higher than the threshold, the facial expression capture result is output as video data; if the confidence level is lower than the threshold, a second facial capture model is switched within the candidate region for detection, and the facial expression capture result is output as video data.
[0036] Specifically, this embodiment addresses the problem that the forward video acquisition unit may lose faces for a long time due to subjects looking down, tilting their heads, or partially obscuring their faces during the acquisition process. It utilizes a lateral acquisition unit to lock the candidate area of the subject's face, thereby expanding the acquisition range of the forward acquisition unit to acquire facial video data and improve the acquisition quality.
[0037] Specifically, the first facial acquisition model is a DNN model, and the second facial acquisition model is a Haar model.
[0038] In this embodiment, the video data acquisition process further includes: evaluating the quality of the acquired video data, the process of which is as follows: Based on the Euler angles of the subject's head, the subject's posture compliance index was calculated, and based on the brightness distribution of the subject's facial area, the illumination compliance index was calculated. The posture compliance index and the illumination compliance index are compared with preset quality thresholds. If both indexes are greater than the preset quality thresholds, video data acquisition is continued. If either index is lower than the preset quality threshold, the subject is guided to adjust their posture until both indexes are greater than the preset quality thresholds, and then video data acquisition is resumed. Based on the statistical analysis and guidance records of indicators during the video data acquisition process, data quality level labels are generated for the video data. The data quality level labels are then associated with and stored with the video data to complete the acquisition of multimodal depressive mood data.
[0039] Specifically, the guidance of the subject is accomplished through a display screen set in front of the subject. When either the posture qualification index or the illumination qualification index is lower than the preset quality threshold, the subject is guided to perform actions such as "please look up at the screen" or "please adjust your position to get brighter light" through the display screen prompts. When the posture and illumination indices are higher than the preset quality threshold for several consecutive frames, the playback of the stimulus material and data collection are automatically resumed, and the intervention event is recorded.
[0040] Specifically, by adjusting and evaluating the quality of video data, the quality control of video data is moved forward to the acquisition stage, significantly improving the effective sample ratio and signal-to-noise ratio of the final dataset.
[0041] In this embodiment, the audio data acquisition process further includes: using a callback function to receive real-time audio data, and using a mutex lock to prevent other operations from writing the audio data during the receiving process.
[0042] Specifically, the purpose of a mutex lock is to prevent other threads from writing to the audio file during the execution of the callback function, ensuring the atomicity of each write operation and the consistency of the data.
[0043] Specifically, if an error occurs during the audio writing process, the system will restore the file writing state to the previous valid state through an error rollback mechanism to avoid file corruption or data loss caused by partial writing failure.
[0044] Specifically, after audio data acquisition is complete, a file integrity check is automatically performed to ensure that all recording data is correctly saved. This effectively prevents audio data fragmentation caused by concurrent writes in a multi-threaded environment, improving the consistency and reproducibility of audio data.
[0045] Specifically, after the video and audio data are collected, a session directory is constructed based on the subject's identity information. The video and audio data that have undergone drift compensation are stored in the session directory according to their data types, thus completing the collection of multimodal depressive mood data.
[0046] Example 3 This embodiment proposes a data acquisition system for multimodal depressive mood data. In this embodiment, the system is used to implement a method for acquiring multimodal depressive mood data. A schematic diagram of the system is shown below. Figure 3 As shown, it includes: The synchronous data acquisition module is used to synchronously start the acquisition of video data and audio data based on a hardware triggering mechanism; the video data includes the subject's facial video data and the subject's gait video data, and the audio data includes the subject's questionnaire recording audio data and the subject's reading audio data. The drift compensation module is used to construct a unified time reference. Based on the preset sampling parameters of video data and audio data and the unified time reference, it calculates the time deviation of video data and audio data, and performs drift compensation on the acquired video data and audio data based on the time deviation. The data storage module is used to collect multimodal depressive mood data by using the drift-compensated video and audio data as the final multimodal depressive mood data.
[0047] Example 4 In this embodiment, a computer device 100 is proposed, which includes a memory 101, a processor 102, and a computer program stored in the memory 101 that can be executed by the processor. The processor 102 executes the computer program to implement a method for collecting multimodal depressive mood data. A schematic diagram of the device is shown below. Figure 4 As shown.
[0048] In this embodiment, a computer storage medium is also proposed, on which a computer program is stored. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform a method for collecting multimodal depressive mood data.
[0049] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, and are not intended to limit the implementation of the present invention. Those skilled in the art can make other variations or modifications based on the above description. It is neither necessary nor possible to exhaustively describe all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A multi-modal depression mood data collection method, the method is applied to a multi-modal depression mood data collection system, characterized in that, The method comprises the following steps: Synchronously starting collection of video data and audio data based on a hardware trigger mechanism; the video data comprises facial video data of a subject and gait video data of the subject, and the audio data comprises questionnaire recording audio data of the subject and reading audio data of the subject; Building a unified time reference, calculating a time deviation of the video data and the audio data based on preset sampling parameters of the video data and the audio data and the unified time reference, and performing drift compensation on the collected video data and audio data based on the time deviation; Taking the video data and the audio data that have been subjected to drift compensation as final multi-modal depression emotion data, and completing collection of the multi-modal depression emotion data.
2. The method for collecting multi-modal depression mood data according to claim 1, wherein, When the collection of the video data and the audio data is synchronously started based on the hardware trigger mechanism, all collection units are synchronously started based on a unified hardware clock, and the collection of the video data and the audio data is started; the collection units comprise a front video collection unit, a side video collection unit and an audio collection unit. The facial video data of the subject is collected by the front video collection unit, the gait video data of the subject is collected by the front video collection unit and the side video collection unit in cooperation, and the audio data is collected by the audio collection unit.
3. The method for collecting multi-modal depression mood data according to claim 2, wherein, The preset sampling parameters include a video target frame rate and an audio sampling rate The process of calculating the time difference between the video data and the audio data is: based on a start timestamp of a forward video capture unit activation , a video frame sequence number , and a video target frame rate calculating a video time , expressed as: based on a start time stamp of the audio acquisition unit at the time of activation , the number of audio samples , and the audio sampling rate to calculate the audio time , the expression is based on the video time and the audio time a time offset between the audio time and the video time is calculated as 。 4. The method for collecting multi-modal depression mood data according to claim 3, wherein, Based on the time deviation, the collected video data and audio data are subjected to drift compensation, and the process is as follows: Determining time deviation between audio time and video time whether the first time difference is less than a preset first time difference threshold, and if so, not performing synchronization of the video data and the audio data; if the time deviation between the audio time and the video time is greater than a preset first time difference threshold and less than a preset second time difference threshold, then adjusting the sampling rate if the time deviation between the audio time and the video time is greater than a preset first time difference threshold and less than a preset second time difference threshold, then adjusting the sampling rate until the time deviation is less than the first time difference threshold less than the first time difference threshold If the time difference between audio time and video time If the time difference is greater than or equal to the second time difference threshold, synchronization is performed based on the deviation direction between the video and audio data. If the video data lags behind the audio data, non-critical frames in the video data are discarded until the time difference is reached. If the time difference is less than the first time difference threshold; if the video data leads the audio data, then perform frame interpolation on the non-key frames in the video data until the time difference is reached. Less than the first time difference threshold.
5. The method for collecting multi-modal depression mood data according to claim 3, wherein, In the video data collection process, the following steps are further included: A candidate region of the face and the gait of the subject is calculated, and the facial video data and the gait video data of the subject are collected based on the candidate region, and the process is as follows: A head two-dimensional coordinate in a side image of the subject is extracted in real time by the side video collection unit, and the head two-dimensional coordinate is inversely calculated into an estimated position of the head in a three-dimensional space in combination with a camera parameter that is pre-calibrated; Based on a spatial relationship matrix of the front video collection unit and the side video collection unit that is pre-calibrated, the head three-dimensional space position is projected to an imaging plane of the front video collection unit, and a candidate region in which the face of the subject appears in the imaging plane of the front video collection unit is calculated; A first face collection model pre-set in the front video collection unit is called in the candidate region to perform facial expression collection, a confidence of a facial expression collection result is obtained, the confidence is compared with a preset confidence threshold, if the confidence is higher than the threshold, the facial expression collection result is output as the video data, and if the confidence is lower than the threshold, a second face collection model is switched to perform detection in the candidate region, and a facial expression collection result is output as the video data.
6. The method for collecting multi-modal depression mood data according to claim 5, wherein, In the video data collection process, the following steps are further included: Based on a head Euler angle of the subject, an attitude qualification index of the subject is calculated, and based on a brightness distribution of a face region of the subject, an illumination qualification index is calculated. The posture qualification index and the illumination qualification index are compared with a preset quality threshold value. If both indexes are greater than the preset quality threshold value, video data collection is maintained. If either index is lower than the preset quality threshold value, the subject is guided to adjust the posture until both indexes are greater than the preset quality threshold value, and then video data collection is resumed. Based on the index statistics and guidance records in the video data collection process, a data quality level label is generated for the video data, and the data quality level label is stored in association with the video data, thereby completing the collection of multi-modal depression emotion data.
7. The method for collecting multi-modal depression mood data according to claim 6, wherein, In the audio data collection process, a callback function is used to receive real-time audio data, and a mutex is used to prevent other operations from writing to the audio data during reception.
8. A multi-modal depression mood data oriented acquisition system, characterized in that, The system is used to implement the method of any one of claims 1-7, comprising: a synchronous data collection module configured to start the collection of video data and audio data based on a hardware triggering mechanism; the video data includes facial video data of the subject and gait video data of the subject, and the audio data includes questionnaire recording audio data of the subject and reading audio data of the subject; a drift compensation module configured to construct a unified time reference, calculate a time deviation of the video data and the audio data based on preset sampling parameters of the video data and the audio data and the unified time reference, and perform drift compensation on the collected video data and audio data based on the time deviation; a data storage module configured to store the video data and the audio data that have been subjected to drift compensation as final multi-modal depression emotion data, thereby completing the collection of multi-modal depression emotion data.
9. A computer device, comprising: The computer device includes a memory, a processor, and a computer program stored on the memory and executable by the processor, the processor executes the computer program to implement the method of any one of claims 1-7.
10. A computer storage medium, characterized in that The computer program includes program instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-7.