Target region motion state detection method, device, equipment and medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-11
AI Technical Summary
[0005]本发明的主要目的在于提供一种目标区域运动状态检测方法、装置、设备及存储介质,旨在解决现有技术在进行视频中目标区域运动检测如嘴部运动检测时,通常依赖人工逐帧标注或直接利用音频时间信息生成视觉标签,容易产生标注成本高、音视频时间偏移以及运动边界识别不准确的技术问题
[0010]有益效果:本发明涉及检测模型技术领域,公开了一种目标区域运动状态检测方法、装置、设备及介质,包括:获取音频数据及同步视频数据,提取声音有效片段和音频帧级能量序列;对声音有效片段进行非对称边界扩展得到目标有效片段,并结合音频帧级能量序列生成视频数据中目标区域的初始软标签序列;提取目标区域的视觉特征序列并进行模型训练得到初始运动检测模型;利用初始运动检测模型生成初始预测序列并确定不确定度序列;根据不确定度序列融合初始预测序列与初始软标签序列得到修正后软标签序列;基于修正后软标签序列进行迭代训练得到目标运动检测模型;对待检测视频数据定位目标区域并提取视觉特征序列,输入目标运动检测模型得到运动状态检测结果。本发明通过结合音频信息生成软标签并利用不确定度序列进行标签融合修正,在减少人工逐帧标注的同时缓解音视频时间偏移带来的边界误差,并通过迭代训练优化模型参数,从而提高视频中目标区域运动检测的准确性。
Smart Images

Figure CN122551240A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of detection model technology, and in particular to a method, apparatus, equipment and medium for detecting the motion state of a target area. Background Technology
[0002] With the development of network communication and video interaction technologies, remote interaction methods based on audio and video data are increasingly being applied to various business scenarios. In these scenarios, the system typically needs to confirm whether the person in the video is actually participating in a voice exchange. One important criterion is whether the mouth area produces corresponding movements during speech. However, existing technologies still have significant shortcomings in mouth movement detection. On the one hand, model training usually relies on manually annotated video data frame by frame, requiring confirmation of the mouth state in each frame, a time-consuming and costly process. On the other hand, some technical solutions rely on audio signals to generate visual labels, but since there is a natural temporal shift between mouth movements before and after speech and the speech signal, directly using audio time information as visual labels can easily introduce temporal boundary errors. Furthermore, mouth movements at the beginning and end of speech typically exhibit a gradual and continuous process, which traditional binary labels cannot reflect, easily leading to unstable predictions by the model at boundary positions, thus affecting the accuracy of mouth movement detection.
[0003] In the fintech sector, financial institutions typically use video interaction to verify customer identity and conduct in processes such as remote account opening, remote identity verification, and online interviews. This process carries the risk of someone else answering questions on behalf of the interviewee; for example, a customer might appear on screen, but the person actually answering questions could be someone else off-screen. To mitigate this risk, systems often detect the movement of the customer's mouth area while speaking to determine the consistency between the audio source and the subject on screen. However, relying on extensive manual frame-by-frame annotation of training data, or directly generating visual labels from audio timestamps, can lead to high annotation costs, inconsistencies between audio and video timing, and inaccurate motion boundary recognition, thus affecting the reliability of mouth movement detection in remote interview scenarios.
[0004] In the healthcare sector, telemedicine services such as remote consultations, remote rehabilitation assessments, and online psychological counseling typically involve real-time communication between doctors and patients via audio and video. In these scenarios, the system also needs to verify that the patient in the video is actually speaking, to avoid discrepancies between the voice source and the subject on screen. However, existing technologies still face similar challenges in building mouth movement detection models to those in the financial sector. These include the need for extensive, frame-by-frame manual annotation of video data, reliance on audio information to generate visual labels leading to temporal boundary shifts, and difficulty in accurately describing the gradual changes in mouth movements at the beginning and end of the movement. These issues, in turn, affect the accuracy of mouth movement recognition during telemedicine interactions. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for detecting the motion state of a target region. This invention aims to solve the technical problems of high annotation costs, audio-video time offset, and inaccurate motion boundary recognition in existing technologies for detecting motion in target regions such as mouth movements in videos, which typically rely on manual frame-by-frame annotation or directly using audio time information to generate visual labels.
[0006] To achieve the above objectives, the present invention provides a method for detecting the motion state of a target area, comprising: Acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data; The effective sound segment is asymmetrically extended to obtain the target effective segment, and an initial soft tag sequence of the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence. Extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain the initial motion detection model; An initial prediction sequence is obtained by predicting the motion state of the visual feature sequence using the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined. The initial prediction sequence and the initial soft label sequence are fused based on the uncertainty sequence to obtain the corrected soft label sequence. The initial motion detection model is iteratively trained based on the corrected soft label sequence and the initial prediction sequence. Training stops when a preset termination condition is met, and the target motion detection model is obtained. Acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0007] Furthermore, to achieve the above objectives, the present invention provides a target area motion state detection device, comprising: The audio and video preprocessing module is used to acquire audio data and video data synchronized with the audio data, and to extract effective sound segments and audio frame-level energy sequences from the audio data. The soft tag generation module is used to perform asymmetric boundary expansion on the effective sound segment to obtain the target effective segment, and generate an initial soft tag sequence for the target region contained in the video data based on the target effective segment and the audio frame-level energy sequence. An initial training module is used to extract the visual feature sequence of the target region in the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain an initial motion detection model. An uncertainty assessment module is used to predict the motion state of the visual feature sequence using the initial motion detection model to obtain an initial prediction sequence and determine the uncertainty sequence of the initial prediction sequence. The label fusion correction module is used to fuse the initial prediction sequence and the initial soft label sequence according to the uncertainty sequence to obtain the corrected soft label sequence; An iterative optimization module is used to iteratively train the initial motion detection model based on the corrected soft label sequence and the initial prediction sequence. Training is stopped when a preset termination training condition is met, and the target motion detection model is obtained. The inference detection module is used to acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a target area motion state detection program stored in the memory and executable on the processor, wherein when the target area motion state detection program is executed by the processor, it implements the steps of the target area motion state detection method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a target region motion state detection program, wherein the target region motion state detection program, when executed by a processor, implements the steps of the target region motion state detection method as described above.
[0010] Beneficial Effects: This invention relates to the field of detection model technology, and discloses a method, apparatus, device, and medium for detecting the motion state of a target region. The method includes: acquiring audio data and synchronized video data; extracting effective sound segments and audio frame-level energy sequences; performing asymmetric boundary expansion on the effective sound segments to obtain target effective segments, and combining the audio frame-level energy sequences to generate an initial soft label sequence for the target region in the video data; extracting the visual feature sequence of the target region and training the model to obtain an initial motion detection model; using the initial motion detection model to generate an initial prediction sequence and determine an uncertainty sequence; fusing the initial prediction sequence and the initial soft label sequence based on the uncertainty sequence to obtain a corrected soft label sequence; iteratively training based on the corrected soft label sequence to obtain a target motion detection model; locating the target region in the video data to be detected and extracting the visual feature sequence, inputting it into the target motion detection model to obtain the motion state detection result. This invention, by combining audio information to generate soft labels and using the uncertainty sequence for label fusion correction, reduces manual frame-by-frame annotation while mitigating boundary errors caused by audio-video time offsets. Furthermore, it optimizes model parameters through iterative training, thereby improving the accuracy of motion detection in target regions of videos. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a target region motion state detection method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the target region motion state detection method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the target area motion state detection device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The target region motion state detection method provided in this embodiment of the invention can be applied to, for example... Figure 1In this application environment, the client communicates with the server via a network. The server can obtain audio data and synchronized video data from the client, extract effective sound segments and audio frame-level energy sequences; perform asymmetric boundary expansion on the effective sound segments to obtain target effective segments, and combine the audio frame-level energy sequences to generate an initial soft label sequence for the target region in the video data; extract the visual feature sequence of the target region and train the model to obtain an initial motion detection model; use the initial motion detection model to generate an initial prediction sequence and determine the uncertainty sequence; fuse the initial prediction sequence and the initial soft label sequence according to the uncertainty sequence to obtain a corrected soft label sequence; perform iterative training based on the corrected soft label sequence to obtain a target motion detection model; locate the target region in the video data to be detected and extract the visual feature sequence, input it into the target motion detection model to obtain the motion state detection result. This invention, by combining audio information to generate soft labels and using uncertainty sequences for label fusion correction, reduces manual frame-by-frame annotation while mitigating boundary errors caused by audio-video time offset, and optimizes model parameters through iterative training, thereby improving the accuracy of motion detection in target regions of the video. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the target area motion state detection method provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the target region motion state detection method proposed in this invention includes the following steps: S10, acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data; In this embodiment, the synchronous acquisition of audio and video data is used to form a multimodal data set on a unified time axis, ensuring that changes in speech signals and changes in image frames maintain a correspondence on the same time coordinate. Audio data represents continuous acoustic signals acquired by a sound acquisition device, containing information about energy changes caused by sound vibrations. Video data represents a sequence of continuous image frames acquired by a camera device, with each frame recording the visual state of a target object at a specific time location. In a remote interactive environment, the audio acquisition device can be from a terminal microphone array or a built-in acquisition unit of the communication terminal, and the video data can be from a terminal camera device or the image acquisition module of a video conferencing system. Synchronization is established through a unified timestamp, for example, by writing a unified timestamp for audio and video frames at the data acquisition end, or by performing time alignment processing on the two types of data at the server end based on streaming media time information. Once the audio frame time and video frame time correspond on a unified time axis, it ensures that the time of speech activity and visual changes in image frames can be subsequently correlated and analyzed.
[0016] A valid audio segment represents the time interval within audio data where speech activity occurs. Audio signals exhibit continuous waveforms on the time axis, and speech activity intervals are typically accompanied by significant energy changes. Therefore, audio frame data can be obtained by segmenting the audio signal into time segments. An audio frame represents a segment of acoustic signal captured within a fixed time window. Each audio frame contains several sampling points; the energy value of that time window is obtained by squaring and summing the amplitudes of the sampling points. The energy value reflects changes in acoustic vibration intensity; when the energy value exceeds the ambient noise baseline, it can be identified as a speech activity interval. When the energy values of multiple consecutive audio frames exceed a threshold, these audio frames can be merged on the time axis to form a speech activity interval. This time interval constitutes a valid audio segment. A valid audio segment includes two time parameters: a start time position and an end time position, used to represent the duration of the speech activity. In a fintech remote interview environment, a valid audio segment typically corresponds to the speech interval during a client's answer to a question. In a healthcare remote consultation environment, a valid audio segment corresponds to the time interval during which the person seeking consultation expresses their thoughts.
[0017] Audio frame-level energy sequences are used to describe the energy variation trend of audio signals along the time axis. An audio frame-level energy sequence consists of the energy values of multiple audio frames arranged chronologically. The calculation of energy values depends on the amplitude variation information of the audio signal. Each audio frame forms several sampling points during the sampling phase. The amplitude of each sampling point is squared and averaged to obtain the energy amplitude. This energy amplitude reflects the acoustic signal intensity within that time window. The energy amplitudes corresponding to multiple time windows are arranged chronologically to form the audio frame-level energy sequence. The audio frame-level energy sequence can not only indicate the presence of speech activity but also describe the changes in speech intensity, such as the gradual increase in energy before speech and the gradual decrease in energy after speech. The audio frame-level energy sequence establishes a correspondence with the video frame time axis through timestamp mapping, so that each video frame position corresponds to an energy value. In this way, changes in the speech signal can establish a temporal correspondence with changes in visual state in video image frames.
[0018] There is a complementary information relationship between effective sound segments and audio frame-level energy sequences. Effective sound segments provide information on the temporal range of speech activity, while audio frame-level energy sequences provide information on speech intensity changes. Effective sound segments can determine the approximate start and end times of speech activity, while audio frame-level energy sequences can reflect the gradual increase or decrease in energy during speech activity. In remote video interaction environments, these two types of information can be used to identify the temporal location of speech expressions and provide a temporal reference for image data analysis. In fintech remote auditing scenarios, the time interval for a customer's voice answering questions can be determined using effective sound segments, and speech intensity changes can be described using audio frame-level energy sequences. In healthcare remote consultation scenarios, the duration and intensity changes of the consultant's speech expressions can also be represented using the above information.
[0019] The identification of effective sound segments can be enhanced by combining acoustic feature analysis. For example, in addition to energy values, spectral distribution features or formant information can be extracted to distinguish speech signals from background noise. Spectral features obtain frequency distribution information through short-time Fourier transform. Speech signals usually exhibit significant energy concentration within a specific frequency range; identifying these frequency distribution features can further confirm the speech activity interval. Audio frame-level energy sequences can be normalized during construction to ensure energy values are within a uniform range under different acquisition environments, thus facilitating subsequent data processing.
[0020] This embodiment, by simultaneously acquiring audio and video data and extracting effective sound segments and audio frame-level energy sequences, can determine the duration of speech activity and obtain speech intensity variation information on a unified timeline. Effective sound segments provide information on speech activity intervals, while audio frame-level energy sequences provide information on energy change trends. Combining these two types of information allows for a more accurate description of the temporal changes in speech expression, thus providing a reliable time reference for subsequent image data analysis and reducing the need for manual frame-by-frame annotation.
[0021] S20, the effective sound segment is asymmetrically extended to obtain the target effective segment, and the initial soft tag sequence of the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence; In this embodiment, the effective sound segment represents the time interval in the audio data where speech activity exists, which is obtained through a speech activity recognition process. Speech activity recognition typically determines the duration of speech based on changes in audio signal energy, spectral distribution, and speech structure characteristics. The effective sound segment includes the start and end times of the speech activity, and the time parameter allows for locating the range of speech expression on a unified time axis. Because there is a physiological time difference between speech production and mouth movements—the lips begin preparing before phonation and remain closed for a period after phonation ends—the time interval corresponding to the speech signal cannot completely cover the actual range of mouth movement changes. To address this time offset problem, asymmetric boundary expansion processing is needed for the effective sound segment to more accurately cover the actual range of mouth movement changes.
[0022] Asymmetric boundary extension refers to adjusting the start and end boundaries of the effective sound segment with different time extension amplitudes. Mouth movements typically occur earlier in the pre-phonation phase, so an extension time is added before the start time of the effective sound segment to cover the time range of pre-phonation mouth preparation. Mouth closure movements in the end phase are typically shorter, so a smaller extension time is added after the end time of the effective sound segment. This different time extension amplitude creates a new time interval, called the target effective segment. The target effective segment covers a wider area on the time axis than the original effective sound segment while maintaining its correlation with the time range of the speech activity.
[0023] Once the target valid segment is established, the set of video frames corresponding to that time interval can be located on the video frame timeline. Video data exists as a continuous sequence of image frames, with each frame containing information about the target object's facial region or other visual regions. Timestamp matching can determine the set of video frame positions corresponding to the target valid segment, thereby identifying the set of image frames where mouth movement may occur within that time interval. The target region represents the visual region of interest in a video frame, such as the mouth region within the facial region. In fintech remote interaction scenarios, the target region typically corresponds to the mouth position within the customer's facial region. In healthcare remote consultation scenarios, the target region can represent the area of the consultant's face region involved in verbal expression.
[0024] The initial soft-label sequence is used to represent the motion probability information of the target region along the time axis. Unlike traditional binary labels, soft labels represent changes in motion probability with continuous numerical values. Soft labels can describe the gradual characteristics of mouth movements over time, such as a gradual transition from no movement to noticeable movement, and then from noticeable movement back to a static state. The generation of the soft-label sequence depends on the temporal information of the target effective segment and the audio frame-level energy sequence. The target effective segment provides information on the temporal range of speech activity, while the audio frame-level energy sequence provides information on speech energy changes. Speech energy is usually correlated with the amplitude of mouth movements; when speech energy increases, mouth movements are usually more pronounced, and when speech energy decreases, mouth movements gradually decrease.
[0025] An audio frame-level energy sequence consists of consecutive audio frame energy values, each energy value corresponding to a position on the timeline. By mapping the audio frame energy values to the video frame timeline, the speech energy information corresponding to each video frame can be obtained. Energy values can be converted into motion probability values within the target effective time segment, for example, by normalizing the energy values to a numerical range between zero and one. Time positions with higher energy values correspond to higher motion probabilities, and time positions with lower energy values correspond to lower motion probabilities. The mapped probability values are arranged in chronological order to form a probability sequence, which is consistent with the video frame timeline.
[0026] Once the probability sequence is aligned with the video frame timeline, probability values can be assigned to the corresponding time positions of the target regions within the video frames. Each video frame contains a probability value for its target region, representing the likelihood of motion occurring in that target region at that time position. The probability values corresponding to consecutive video frames form a time series, which is the initial soft label sequence. This initial soft label sequence maintains the same order as the video frames on the timeline and describes the changes in target region motion over time. In this way, labeled data for training visual recognition models can be generated without manual frame-by-frame annotation.
[0027] The initial soft-label sequence can also reflect changes in motion intensity during speech expression. In the pre-vocalization stage, the energy value gradually increases, corresponding to a rising probability value, indicating a gradual transition of mouth movement from stillness to noticeable motion. In the final stage of speech, the energy value gradually decreases, corresponding to a decreasing probability value, indicating a gradual cessation of mouth movement. This continuous probability representation avoids the abrupt changes at boundary positions inherent in traditional hard labels, making the label data more consistent with real-world motion changes.
[0028] This embodiment generates an initial soft-label sequence by performing asymmetric boundary expansion on effective sound segments and combining it with audio frame-level energy sequences. This allows the speech activity time interval to more accurately cover the range of mouth movement changes. Asymmetric temporal expansion can compensate for the temporal difference between speech signals and visual motion, and audio energy information can reflect the intensity changes during speech expression. Representing the motion changes of the target region in a continuous probabilistic form reduces the need for manual frame-by-frame annotation and mitigates boundary position recognition errors, making the generated label data closer to the actual motion change patterns.
[0029] S30, extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model and obtain the initial motion detection model; In this embodiment, the target region in the video data represents the visual region in the image frame that carries speech-related motion information. The target region can be the mouth region within a facial area, or a local region that represents vocalization in other remote interaction scenarios. The purpose of setting a target region is to exclude irrelevant background, clothing texture, head contours, and other information from the main processing scope of the entire image frame, so that the subsequently extracted visual information focuses on reflecting speech-related motion changes. The video data itself consists of image frames arranged in chronological order, with each frame corresponding to a time position. The target region forms a local image sequence in consecutive image frames. The visual feature sequence is a continuous feature representation obtained by visually encoding these local image sequences. Each feature vector in the visual feature sequence corresponds to a time position, and multiple feature vectors arranged in chronological order form a temporal visual representation used to characterize the morphological, textural, and motion changes of the target region over continuous time.
[0030] In the visual feature sequence, "visual" refers to the source of the image signal, "feature" refers to the compact representation extracted from the image, and "sequence" refers to the continuous organization of these representations in the time dimension. After the image signal enters the processing flow, it is not directly trained using the raw pixels, but is transformed into a numerical representation more suitable for recognizing motion states through a multi-layer feature extraction structure. Visual feature extraction can be accomplished using convolutional coding structures or attention coding structures based on image patch partitioning. In a convolutional coding structure, the input layer receives the local image corresponding to the target region. Shallow convolutional units extract edge, brightness transition, local texture, and contour information; mid-level convolutional units extract combined information such as mouth opening and closing, lip line changes, tooth alignment, and changes in local facial shadows; and deep convolutional units form an abstract representation related to the motion state. If an attention coding structure is used, the local image can first be divided into multiple image patches, then the image patches can be mapped to a unified dimension through an embedding layer. Afterward, multi-head attention units extract the spatial dependencies between image patches, and feature reconstruction is completed through feedforward units. Regardless of the structure used, the output forms a frame-level visual representation corresponding to a single image frame.
[0031] The formation of visual feature sequences involves not only spatial information extraction but also temporal information modeling. A single frame image can reflect the static state at a specific moment, but motion detection of a target region depends on the changing relationships between consecutive moments. Therefore, it is necessary to further organize the frame-level visual representation into a temporally correlated representation. Temporally correlated structures can be achieved through temporal convolutional units, recurrent units, or temporal attention units. Temporal convolutional units extract change patterns between adjacent time positions by setting continuous convolutional windows on the time axis, used to characterize the local continuity of mouth opening and closing, continuous vocalization, and closure processes. Recurrent units record state transfer between preceding and subsequent time positions through gating structures, used to characterize continuous motion trends in speech expression. Temporal attention units establish dependencies across the entire time range, identifying state relationships between image frames that are far apart. After temporal modeling, the original frame-level representation containing only spatial information is transformed into a visual feature sequence that combines both spatial and temporal information.
[0032] The initial soft-label sequence provides supervision information for the visual feature sequence. Each label value in the initial soft-label sequence corresponds to a time position, and the magnitude of the value reflects the probability of motion occurring in the target region at that time position. Unlike binary labels, soft labels can express continuous transition states from stillness to motion and from motion to stillness. The goal of model training is to learn the motion representation of image regions through visual feature sequences, and then constrain the learning direction through the initial soft-label sequence, so that the model output gradually approaches the motion probability distribution characterized by the labels. In model training, "model" refers to the parameterized structure that undertakes the task of motion recognition of the target region, "training" refers to the process of adjusting parameters using sample data, and "initial motion detection model" refers to the detection model obtained after completing the current stage of training but not yet entering the subsequent correction and iteration stage. Structurally, the model includes at least a visual input layer, a spatial encoding layer, a temporal modeling layer, and a state output layer. The visual input layer is responsible for receiving local image sequences of the target region, the spatial encoding layer is responsible for extracting spatial representations, the temporal modeling layer is responsible for extracting temporal change representations, and the state output layer is responsible for mapping temporal features to motion state probabilities. Layers are connected via tensor transfer structures. The output of the spatial coding layer is connected to the input of the temporal modeling layer, and the output of the temporal modeling layer is connected to the input of the state output layer, forming a continuous data transformation path.
[0033] The model training process revolves around the correspondence between the visual feature sequence and the initial soft label sequence. When training data enters the system, a one-to-one correspondence between image frame positions and soft label positions needs to be established, ensuring that each visual feature vector can find a label value at the same time position. After establishing the correspondence, the visual feature sequence is fed as input to the state output layer, which outputs a set of predicted values corresponding to the time positions. The predicted values are compared with the initial soft label sequence to measure the error, and the error measurement result is then backpropagated to the parameters of each layer, driving parameter updates. Error measurement can use either cross-entropy loss or squared difference loss. If the label values are within a continuous interval, cross-entropy loss is used to measure the difference in probability distribution, while squared difference loss is used to measure numerical deviation. To avoid drastic oscillations in predicted values between adjacent time positions, a smoothing constraint term can be added to the error measurement. This smoothing constraint term suppresses excessive fluctuations by comparing the magnitude of the difference between predicted values at adjacent time positions. Parameter updates can use adaptive gradient descent or momentum gradient update. During training, training parameters such as batch size, learning rate, time window length, and feature dimension need to be set. Batch size controls the number of samples involved in a single update, learning rate controls the magnitude of parameter adjustments each time, time window length determines the range of image frames in a single input, and feature dimension determines the numerical length of the encoded result. If the time window is too short, it's difficult to reflect continuous motion trends; if the time window is too long, it increases training costs. If the learning rate is too large, parameter updates are prone to instability; if the learning rate is too small, the training process converges slowly. During training, the learning rate can be adjusted based on changes in the loss value, or training can be terminated based on changes in the output results on the validation data.
[0034] In the fintech business, the training process for visual feature sequences and initial soft-label sequences is typically deployed in data processing platforms for business scenarios such as remote face-to-face interviews, remote account opening, remote credit approval, and remote dual-recording quality inspection. Input data can come from online video interaction data of customers, and the target region can be set as the mouth area of the face. The visual feature sequence corresponds to the continuous changes in the mouth image when the customer answers audit questions, and the initial soft-label sequence corresponds to the motion probability information generated by audio-driven processes. The trained initial motion detection model is used to identify whether there is synchronous motion in the mouth area when the customer speaks. If the main subject in the video does not produce corresponding motion but the audio continues, it can provide input for subsequent consistency judgments. In the healthcare business, input data can come from interactive video data such as remote psychological counseling, remote rehabilitation guidance, and online consultation and care audits. The target region can also be set as the mouth area of the face. The visual feature sequence is used to describe the local motion changes of the consultation subject during continuous expression, and the initial soft-label sequence is used to provide training constraints, enabling the output model to recognize the motion state of the target region during speech expression.
[0035] This embodiment extracts visual feature sequences from the target region and trains them in conjunction with an initial soft-label sequence, enabling the continuous representation of motion change information in the image region over time. The spatial coding structure extracts local morphological and texture changes in the target region, the temporal modeling structure extracts motion relationship between consecutive image frames, and the initial soft-label sequence provides probabilistic constraints on the model output. The resulting initial motion detection model learns the correspondence between the motion state of the target region and image changes from video data, giving the subsequent recognition process trainable visual judgment capabilities.
[0036] S40, the initial prediction sequence is obtained by predicting the motion state of the visual feature sequence through the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined. In this embodiment, the initial motion detection model represents a parameterized visual recognition structure formed under the constraints of existing training data. This structure receives a sequence of visual features and outputs a motion state prediction result at the corresponding time position. The visual feature sequence comes from the visual encoding results of image frames of the target region in the video data. Each time position corresponds to a visual feature vector, and multiple visual feature vectors are arranged in chronological order to form the sequence input. The visual representation in the visual feature sequence comes from morphological changes, contour changes, texture changes, and local brightness changes in the image space. These changes can reflect the motion behavior of the target region in continuous time. After the visual feature sequence is input into the initial motion detection model, state inference is completed through the internal structure of the model, and the inference results are arranged in chronological order to form the initial prediction sequence.
[0037] Motion state prediction involves determining whether motion has occurred in a target region at each time point. The prediction process is completed by multiple processing layers in the initial motion detection model. The model structure typically includes a visual input layer, a feature mapping layer, a temporal representation layer, and a state output layer. The visual input layer receives the visual feature sequence and performs dimensional alignment, ensuring each visual feature vector enters a unified feature space. The feature mapping layer enhances the spatial representation of the input features through parameter matrix mapping or convolutional units, making the original visual representation more prominent in terms of local motion changes in the target region. The temporal representation layer models the temporal relationships between feature vectors at consecutive time points, thereby extracting dynamic change information between consecutive frames. The state output layer converts the temporal representation results into motion state distribution values or probability values. The output results are arranged in chronological order to form the initial prediction sequence, with each predicted value corresponding to a time point in the video frame sequence.
[0038] The initial prediction sequence represents the set of motion state judgment results generated by the model under the current parameter conditions. Predicted values can be represented as continuous probability values or as multi-class distribution values. When the predicted values are expressed as probabilities, each time position corresponds to a value representing the probability that the target region is in motion at that time position. The closer the probability value is to the high value range, the higher the probability of motion in the target region; the closer the probability value is to the low value range, the lower the probability of motion in the target region. When the predicted values are expressed as distributions, each time position corresponds to a set of distribution values, and each distribution value represents the relative probability of different state categories at that time position. The initial prediction sequence is consistent with the visual feature sequence in the time dimension, ensuring that each visual feature vector can find a corresponding prediction result.
[0039] An uncertainty sequence describes the stability of prediction results at each time point in the initial prediction sequence. During motion state prediction, the model's judgments at certain time points may be relatively clear, while judgments at others may have ambiguous areas. The uncertainty sequence reflects the reliability of the model's judgments by quantifying the dispersion or concentration of the prediction results. If the prediction distribution is concentrated around a single state, the corresponding uncertainty is low; if the prediction distribution is dispersed across multiple states, the corresponding uncertainty is high. By analyzing the prediction distribution at each time point in the initial prediction sequence, quantified values reflecting the prediction stability can be extracted. These quantified values, arranged in chronological order, form the uncertainty sequence.
[0040] Uncertainty can be quantified by calculating the entropy of the probability distribution. The entropy of the prediction distribution reflects the degree of uncertainty; a higher entropy indicates a more uniform distribution and a more ambiguous judgment of the state by the model, while a lower entropy indicates a more concentrated distribution and a clearer judgment of the state by the model. Another quantification method is to extract the magnitude of the difference between the highest and second-highest probability values in the prediction distribution. A smaller difference indicates that multiple states have similar probabilities, corresponding to a higher degree of uncertainty; a larger difference indicates that a certain state has a significantly higher probability, corresponding to a lower degree of uncertainty. Regardless of the quantification method used, the obtained uncertainty values are consistent with the time position in the initial prediction sequence and are arranged in chronological order to form an uncertainty sequence.
[0041] The uncertainty sequence maintains the same length as the initial prediction sequence in structure, with each time position corresponding to an uncertainty value. The length of the visual feature sequence is determined by the number of video frames; therefore, the lengths of both the initial prediction sequence and the uncertainty sequence are consistent with the video frame time positions. The correspondence between time positions can be maintained through frame indices or timestamps, aligning the visual feature vectors, predicted values, and uncertainty values on the same time axis. This alignment allows subsequent processing stages to analyze the prediction stability at specific time positions.
[0042] In the fintech sector, video data typically originates from remote face-to-face interviews, remote account openings, or remote identity verification processes. The target region can be set as a local facial area in the video frame, and the visual feature sequence represents the customer's local visual changes over continuous time. The initial motion detection model generates motion state prediction results based on the visual feature sequence, forming an initial prediction sequence. The uncertainty sequence reflects the stability of the model's judgments at each time point; when the prediction results are concentrated, it indicates that the motion state of the customer's facial area is relatively clear, while when the prediction results are scattered, it indicates that there is significant uncertainty in the motion judgment at that time point. In the healthcare sector, video data can originate from remote health consultations, remote rehabilitation guidance, or online psychological communication scenarios. The target region can also be set as a local facial area, and the visual feature sequence reflects the visual changes in the local area during the consultation process. The initial prediction sequence output by the model represents the motion state at continuous time points, and the uncertainty sequence represents the reliability of the prediction results at each time point.
[0043] This embodiment uses an initial motion detection model to predict the motion state of a visual feature sequence and determine an uncertainty sequence. It can simultaneously obtain the motion state judgment result of the target area and the corresponding stability of the judgment in the time dimension. The visual feature sequence provides image change information, the model structure extracts the dynamic relationship between continuous time positions, the initial prediction sequence reflects the motion probability distribution at each time position, and the uncertainty sequence reflects the concentration of the prediction distribution. In this way, motion state changes can be identified in video sequences and the reliability of the prediction results can be quantified, enabling subsequent processing to perform further analysis based on the prediction stability.
[0044] S50, the initial prediction sequence and the initial soft label sequence are fused according to the uncertainty sequence to obtain the corrected soft label sequence; In this embodiment, the uncertainty sequence is used to characterize the stability of the initial prediction sequence at various time points. The initial prediction sequence comes from the inference results of the motion state recognition structure on the visual feature sequence, and the initial soft label sequence comes from the label expression after mapping audio-related information to the time range. The two types of sequences differ in their sources and information attributes. The initial prediction sequence reflects the visual side's judgment result on the motion state of the target area, while the initial soft label sequence reflects the motion probability expression guided by the audio side. The role of the uncertainty sequence is to measure the credibility of the visual side's judgment at various time points, thereby providing a basis for the fusion of the two types of sequences. When the uncertainty at a time point is low, it indicates that the visual side's judgment has strong concentration and high stability, and this time point is more suitable for retaining the information in the initial prediction sequence. When the uncertainty is high, it indicates that the visual side's judgment has obvious ambiguity or dispersion, and this time point is more suitable for increasing the participation of the initial soft label sequence. Uncertainty in the middle range indicates that both the visual side's judgment and the audio side's label have reference significance, and therefore are more suitable for weighted synthesis.
[0045] The foundation of fusion processing lies in time alignment. The initial prediction sequence, initial soft label sequence, and uncertainty sequence need to maintain a correspondence on the same time axis so that predicted values, label values, and uncertainty values at the same time position can participate in the fusion. Time alignment can be achieved based on video frame indices or unified timestamps. When video data is organized as consecutive image frames, each frame corresponds to a fixed time position. Only after the predicted values in the initial prediction sequence, the label values in the initial soft label sequence, and the quantized values in the uncertainty sequence are mapped to the same frame position can position-by-position fusion be performed. If there are differences in sampling density along the time axis, the sequence length can be unified through interpolation mapping, time window alignment, or neighboring position matching.
[0046] The initial prediction sequence in the fusion process represents the visual recognition result. The advantage of visual results lies in their direct origin from changes in the target region image, reflecting local deformation, texture changes, and opening / closing dynamics. The initial soft label sequence represents the audio guidance result. The advantage of audio results is their ability to provide continuous temporal reference information on vocal activity, particularly suitable for time locations with clearly defined vocal duration intervals. The uncertainty sequence serves as an adjustment basis, adjusting the proportion of visual and audio information through quantization values at different time locations. This fusion process is not a simple replacement, but rather assigns different weights based on the stability of the judgment at different time locations, ensuring that the fusion result retains the clearly expressed state from the visual side while also absorbing the audio guidance information at ambiguous boundaries.
[0047] The corrected soft-label sequence represents the label representation after fusion processing. Each label value in the corrected soft-label sequence corresponds to a time position, reflecting the overall probability of motion occurring in the target region at that time position. Compared to the initial soft-label sequence, the corrected soft-label sequence no longer relies solely on audio-side guidance information but combines visual prediction results and stability metrics. Compared to the initial prediction sequence, the corrected soft-label sequence no longer completely depends on the output of the current visual model but avoids directly using fuzzy predictions as labels through uncertainty constraints. The resulting label sequence retains temporal continuity, reduces label jumps at boundary positions, and maintains high consistency within significant motion regions.
[0048] The fusion process can be implemented using a segmented rule approach. In this approach, the uncertainty value is divided into low, medium, and high intervals. The low interval indicates a high concentration of model judgments, allowing the predicted values from the initial prediction sequence to be directly selected as the correction label for the current position. The high interval indicates a high degree of ambiguity in model judgments, allowing the label values from the initial soft label sequence to be selected as the correction label for the current position. The medium interval indicates that both visual judgments and audio labels are meaningful, allowing the predicted values from the initial prediction sequence and the label values from the initial soft label sequence to be combined proportionally to form the correction label for the current position. The proportionality coefficient can be obtained based on the relative relationship between the uncertainty value and the threshold position, or it can be obtained through a preset weight table. Through this segmented rule approach, the fusion strategy at different time positions can automatically switch according to changes in uncertainty.
[0049] Fusion processing can also be implemented using continuous weight mapping. In continuous weight mapping, the uncertainty value is converted into continuously changing visual and audio weights. The visual weight increases as uncertainty decreases, while the audio weight increases as uncertainty increases. The corrected label value at each time point is composed of the weighted sum of the predicted value and the label value. Continuous weight mapping can avoid sudden switching of segmentation rules near the threshold, making the corrected soft label sequence smoother on the time axis. The weight mapping function can be linear or non-linear compression, making the weight changes in the intermediate intervals more consistent with the distribution of business data.
[0050] Threshold or weight settings in the fusion process need to correspond to the data characteristics. In the fintech business, customer answers in remote face-to-face video interviews are typically short and fast-paced, with more concentrated movement changes in the target area. Threshold settings can be biased towards enhancing the visual weights in low-uncertainty intervals, ensuring that clear motion segments maintain a strong visual judgment function. In the healthcare business, voice expressions in remote consultation scenarios may be longer, with slower local movement changes. Uncertainty distributions tend to form wider intermediate intervals at boundary positions. In this case, the fusion interval can be appropriately expanded to make label correction smoother. Regardless of the setting method used, the input data consists of the initial prediction sequence, the initial soft label sequence, and the uncertainty sequence, while the output data consists of the corrected soft label sequence. The output length is consistent with the number of input time positions.
[0051] The formation of the corrected soft tag sequence depends not only on the information at the current location but also on temporal continuity. If the uncertainty value at a certain time location guides the selection of the initial prediction sequence, and adjacent time locations clearly depend on the initial soft tag sequence, direct switching can easily cause local breaks. To mitigate this break, temporal smoothing constraints can be added after the fusion process. For example, the corrected tag values at adjacent time locations can be continuously corrected to keep the variation amplitude between consecutive frames within a reasonable range. Temporal smoothing does not change the dominant basis of the fusion process but rather limits local fluctuations in the temporal dimension after the fusion result is formed, making the corrected soft tag sequence more consistent with the continuous change pattern of the target area's actual movement from stillness to movement and back to stillness.
[0052] This implementation fuses the initial prediction sequence and the initial soft label sequence based on the uncertainty sequence. This allows for the differentiation between regions with clear visual judgments and regions with ambiguous visual judgments at different time points, enabling the corrected soft label sequence to simultaneously absorb visual judgment information and audio guidance information. Time alignment allows the three sequences to participate in processing at the same location, and segmentation rules or continuous weights correspond to different fusion methods for different levels of stability. As a result, the corrected soft label sequence can reduce the bias caused by relying solely on visual predictions or solely on audio labels, and maintain more continuous label changes at boundary positions.
[0053] S60, the initial motion detection model is iteratively trained based on the corrected soft label sequence and the initial prediction sequence. Training is stopped when a preset termination training condition is met, and the target motion detection model is obtained. In this embodiment, the corrected soft-label sequence represents a set of continuous label data corresponding to video frame positions in the time dimension. Each label value describes the motion probability or degree of motion of the target region at the corresponding time position. The initial prediction sequence represents the time-series prediction results generated by the initial motion detection model based on the visual feature sequence. Each prediction value corresponds to the motion state judgment at the same time position. The initial motion detection model represents the visual recognition structure formed after the previous training stage, which can receive the visual feature sequence of the target region and output the motion state expression in the time dimension. Iterative training based on the corrected soft-label sequence and the initial prediction sequence means continuing to update the parameters multiple times based on the current model parameters, so that the model output gradually approaches the corrected label expression. A preset termination training condition is used to determine whether the training process has ended. When the training reaches the preset constraint, the parameter update stops and the current model state is retained, resulting in the target motion detection model.
[0054] Iterative training relies on temporal alignment. Each label value in the corrected soft-label sequence corresponds to the same temporal position as each predicted value in the initial prediction sequence. This temporal position is typically determined by the video frame index or a uniform timestamp. Temporal alignment ensures that the supervised data and prediction results are compared at the same location, thus forming an effective bias metric. If the temporal positions are inconsistent, the label values and predicted values cannot accurately reflect the motion state in the same video frame, and the direction of model parameter updates will be offset. Therefore, temporal alignment is a fundamental condition in the training process.
[0055] The initial motion detection model structurally comprises a visual input module, a temporal feature modeling module, and a state output module. The visual input module receives the visual feature sequence of the target region and maps the feature vector at each time point to a unified feature dimension. The temporal feature modeling module analyzes the dynamic relationships between consecutive time points, enabling the model to identify the trend of the target region changing from static to dynamic and back to static. This module can employ temporal convolutional units, gated recurrent units, or temporal attention units. Temporal convolutional units extract change information between neighboring locations through time windows, gated recurrent units store temporal dependencies through internal state storage, and temporal attention units establish associations between different time points through weight allocation mechanisms. The state output module maps the temporal feature representation to motion state probabilities or distribution values, generating a prediction result for each time point. The prediction results at multiple time points are arranged chronologically to form the initial prediction sequence.
[0056] During iterative training, it is necessary to establish a deviation metric between the predicted results and the corrected labels. This deviation metric measures the degree of difference between the model output and the corrected soft-label sequence. If the corrected labels are expressed as continuous probabilities, the deviation metric can use the squared difference function, calculating the error value by comparing the predicted value with the label value. If the model output is expressed as a state distribution, the deviation metric can use the cross-entropy function, comparing the error value by comparing the predicted distribution with the label distribution. A larger error value indicates a greater deviation between the model's prediction at that time point and the corrected label; a smaller error value indicates that the model output is closer to the label representation.
[0057] Maintaining the continuity of time-series output is also crucial during training. The movement of the target region typically exhibits continuous change over time, rather than abrupt changes; therefore, the model output should maintain a smooth transition between adjacent time points. Temporal continuity constraints can be achieved by constructing a metric for the difference between outputs at adjacent time points. Limiting the difference in predicted values between adjacent time points prevents the model from experiencing drastic oscillations during training. The bias metric and the temporal continuity constraint together constitute the training loss expression. This loss expression guides parameter updates, ensuring that the model output closely approximates the corrected label expression while maintaining a reasonable temporal trend.
[0058] The parameter update process is achieved through backpropagation. Parameters from the visual input module, temporal feature modeling module, and state output module all participate in the update. During training, training parameters such as batch size, learning rate, feature dimension, and time window length need to be set. Batch size represents the number of video segments used in each round of parameter updates. The learning rate represents the magnitude of parameter updates and is used to control the model's convergence speed. Feature dimension represents the length of the visual feature vector, used to ensure dimensionality consistency between the visual input module and the temporal modeling module. The time window length represents the number of consecutive temporal positions input to the model in one training iteration, used to ensure the model can learn continuous motion changes. During training, multiple rounds of parameter updates gradually bring the model output closer to the motion change pattern described by the corrected soft-label sequence.
[0059] Preset training termination conditions are used to control when training stops. The termination condition can be set based on the number of training epochs, stopping training when the preset upper limit is reached. Alternatively, the termination condition can be set based on the magnitude of loss change, stopping training when the loss change is below a preset threshold for several consecutive training epochs. A termination condition can also be set based on the recognition error on the validation data, stopping training when the validation error enters a preset range. After the termination condition is met, the current model parameters are retained and used as the target motion detection model. The target motion detection model maintains the same structure as the initial motion detection model, but its parameters have been iteratively updated to make the model output closer to the corrected label representation.
[0060] This embodiment iteratively trains the initial motion detection model based on the corrected soft-label sequence and the initial prediction sequence, allowing the model output to gradually approach the corrected label representation. The corrected soft-label sequence provides temporal supervision information, the initial prediction sequence reflects the current model judgment state, temporal alignment ensures that the label and prediction result are compared at the same position, and parameter updates enable the model to gradually reduce prediction errors while maintaining continuous temporal change. The target motion detection model obtained after multiple rounds of training can more accurately represent the temporal variation characteristics of the target region's motion.
[0061] S70, acquire the video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0062] In this embodiment, the video data to be detected represents the video input data entering the recognition stage. The data source can be a real-time video stream, an offline video file, or a collection of image frames cached by a remote interactive platform. The video data to be detected is consistent with the video data used in the training stage in terms of data type, both consisting of image frames arranged in chronological order, but they differ in purpose. The video data in the training stage is used to form model parameters, while the video data to be detected is used to call the target motion detection model to perform actual recognition. The video data to be detected retains image change information of the target object over continuous time, including changes in facial contours, local textures, brightness, and the opening and closing of the target area. To ensure that the target motion detection model can stably receive input, the video data to be detected can undergo frame rate unification, resolution unification, brightness normalization, and chronological order verification before entering the recognition process. Frame rate unification is used to keep the time interval between consecutive image frames stable; resolution unification is used to ensure that the image size meets the model input requirements; brightness normalization is used to reduce the differences in image representation caused by different acquisition devices and different ambient lighting; and chronological order verification is used to ensure that the video frame arrangement is consistent with the actual acquisition order.
[0063] The localization operation is used to determine the spatial location of the target region to be detected from the video data to be detected. The meaning of the localization operation is to identify local regions directly related to motion state judgment from the entire frame image, narrowing the input data range from the entire frame image to the region with the main recognition value. The target region to be detected represents the image region that needs to be analyzed in the recognition stage. It can be the mouth region in the facial local region, or other local regions associated with vocalization. The localization operation can be completed through a keypoint localization structure, which determines the boundary of the local region by detecting facial keypoints; it can also be completed through a bounding box regression structure, which determines the range of the target region to be detected by outputting the coordinates of a rectangular region; or it can be completed through a segmentation structure, which extracts local image content through pixel-level region masks. In the keypoint localization structure, the input is the video frame image, and the output is the coordinates of multiple keypoints. The position of the target region to be detected is calculated using the geometric relationship between the keypoints. In the bounding box regression structure, the input is the entire frame image, and the output is the local region boundary parameters. The target region is cropped from the original image based on the boundary parameters. In the segmentation structure, the input is the entire frame image, and the output is a local region mask. The mask is used to preserve the target region and suppress the background region in the image. The result of the localization operation is a set of local images over a continuous time period, with each frame of local image corresponding to a time position in the video data to be detected.
[0064] The process of extracting the visual feature sequence of the target region involves visually encoding continuous local images to form a temporally ordered feature representation. In the visual feature sequence, "to be detected" indicates that the sequence originates from the actual recognition stage rather than the training stage; "visual features" indicates that the features are derived from the extraction of spatial information from the image; and "sequence" indicates that multiple feature vectors are arranged sequentially along the time axis. After the target region enters the encoding process, it is not directly input as raw pixels but is transformed into a numerical feature representation through a visual encoding structure. This visual encoding structure can employ a convolutional extraction module or an image patch embedding and attention extraction module. In the convolutional extraction module, the input layer receives the local image; shallow convolutional units extract edge and texture changes, mid-level convolutional units extract local deformation and contour changes, and deep convolutional units form an abstract motion representation. In the image patch embedding and attention extraction module, the input image is divided into multiple local image patches. Each image patch is mapped to a unified dimension and then enters an attention unit, which forms a high-order feature representation based on the correlation between the image patches. Regardless of the encoding method used, the output is a feature vector corresponding to the time position. Feature vectors at multiple time locations are arranged in the order of video frames to form a sequence of visual features to be detected.
[0065] The target motion detection model represents a parameterized recognition structure that has been trained and meets the recognition requirements. The input to the target motion detection model is a sequence of visual features to be detected, and the output is the motion state detection result of the target region at each time position. The model internally includes at least an input module, a temporal analysis module, and a state output module. The input module receives the sequence of visual features to be detected and performs dimensionality matching, ensuring that feature representations at different time positions enter a unified feature space. The temporal analysis module extracts the change relationships between consecutive time positions, enabling the model to identify opening and closing changes, continuous motion changes, and transition changes of local regions in the time dimension. The temporal analysis module can be composed of temporal convolutional units, gated recurrent units, or temporal attention units. The state output module maps the temporal analysis results to motion state values; the output can be continuous probability values or state category distribution values. There is a feature tensor transfer relationship between the input module and the temporal analysis module, and a state representation mapping relationship between the temporal analysis module and the state output module. If the model adopts a gated recurrent structure, the gated recurrent unit includes an input gate, a state update gate, and a hidden state generation unit, and these units are connected through a parameter matrix. If the model adopts a temporal attention structure, the attention unit includes a query mapping layer, a key mapping layer, and a value mapping layer, and these layers form a temporal dependency representation through attention weight calculation. During the model inference phase, parameter updates are no longer performed; only forward operations are performed to convert the sequence of visual features to be detected into motion state detection results.
[0066] The motion state detection result represents the time-dimensional judgment output of the target motion detection model for the target region to be detected. The motion state detection result can be a motion probability sequence corresponding to a time position, or a state category sequence corresponding to a time position. If the output is a probability sequence, each time position corresponds to a continuous value, the magnitude of which reflects the probability of motion occurring at that time position. If the output is a state category sequence, each time position corresponds to a discrete judgment result, used to distinguish between stationary and moving states. The detection result can be presented as a frame-by-frame output or as an interval-level output through temporal aggregation. Frame-by-frame output is suitable for scenarios requiring fine-grained recognition, while interval-level output is suitable for scenarios requiring overall judgment. The motion state detection result is consistent with the time position in the video data to be detected, ensuring that each video frame corresponds to a recognition result. This provides directly usable recognition information for subsequent business systems, such as identifying when a target region in an interactive video enters or ends a motion state, and whether the motion state exists continuously within certain time intervals.
[0067] In fintech business environments, the video data to be detected typically comes from scenarios such as remote face-to-face interviews, remote account opening, remote credit approval, and remote dual-recording verification. The target area to be detected can be set as a local area of the customer's face, and the visual feature sequence to be detected represents the local image changes of the customer during the question-and-answer process. The target motion detection model takes the visual feature sequence to be detected as input and outputs the motion state detection results of the local area of the customer over continuous time. The input data is set as the local image sequence after localization or the corresponding feature vector sequence, and the output data is set as frame-by-frame motion probability or frame-by-frame state category. In healthcare business environments, the video data to be detected can come from interactive scenarios such as remote psychological communication, remote rehabilitation guidance, and online health consultation. The target area to be detected can also be set as a local area of the face, and the visual feature sequence to be detected reflects the continuous local changes of the consultation subject during the expression process. The target motion detection model outputs the motion state detection results over continuous time, providing a recognition basis for behavioral consistency analysis during remote communication.
[0068] This embodiment acquires the video data to be detected, performs a localization operation to extract the target region to be detected, extracts the visual feature sequence to be detected, and inputs it into the target motion detection model. This allows the recognition process to directly focus on the target region to be detected. The localization operation narrows the scope of image analysis, and the visual feature sequence to be detected provides a continuous temporal representation of local visual changes. The target motion detection model performs temporal analysis on these changes and outputs motion state detection results. Therefore, it is possible to form a target region motion judgment result corresponding to the time and location in the video data to be detected, ensuring that the recognition output is consistent with the actual video content.
[0069] In one embodiment, step S10 above includes: S101, Receive multimedia data stream, perform modal separation operation on the multimedia data stream, and obtain audio data and video data synchronized with the audio data; S102, Perform a time-dimension sliding segmentation operation on the audio data to obtain an audio data frame sequence; S103, Perform a sound boundary localization operation on the audio data frame sequence using the endpoint detection model to generate valid sound segments; S104, extract the short-time energy amplitude of each frame of audio data in the audio data frame sequence, and normalize the short-time energy amplitude to obtain the audio frame-level energy sequence of the audio data.
[0070] In this embodiment, a multimedia data stream represents a continuous data set carrying audio and video information in a unified transmission channel. This data set can originate from a remote interactive terminal, video conferencing platform, online business processing platform, or mobile communication terminal. The data content includes voice sampling data, video frame data, and time stamp information corresponding to the acquisition time. The purpose of receiving the multimedia data stream is to preserve the temporal correspondence between audio and video information, enabling subsequent audio and video data to be matched on the same timeline. The receiving process is typically completed by a data access module, which includes a streaming media receiving unit, a buffering unit, and a time stamp parsing unit. The streaming media receiving unit continuously receives incoming data packets, the buffering unit temporarily stores data content in chronological order, and the time stamp parsing unit reads the time information from the data packets and establishes a unified time sequence. If the business scenario is in a fintech remote interview environment, the received multimedia data stream typically originates from a real-time video conversation between the client and the interviewer; if the business scenario is in a healthcare remote consultation environment, the received multimedia data stream typically originates from remote audio and video interaction between the client and the server.
[0071] Modality separation is used to extract audio and video components from the same multimedia data stream. The result of modality separation includes audio data and video data synchronized with the audio data. Audio data represents continuous acoustic signals captured and encoded by a microphone, while video data represents a sequence of continuous image frames captured and encoded by a camera. Modality separation can be performed through encapsulation format parsing or protocol field parsing. If the multimedia data stream is stored in a containerized form, the encapsulation parsing unit can extract the audio and video tracks according to the stream type identifier; if the multimedia data stream is carried by a real-time transport protocol, the protocol parsing unit can split the audio and video packets according to the payload type identifier. The separated audio and video data still need to retain their temporal correspondence. A common method is to synchronously record the timestamps of each audio segment and each video frame during the separation process, and then construct a synchronization mapping table based on the timestamps. "Synchronization" in synchronized video data means that the video frames maintain a temporal correspondence with the audio data, not that they require the same sampling frequency. Audio data typically records continuous waveforms at a high sampling rate, while video data typically records discrete image frames at a fixed frame rate; the synchronization relationship is maintained through timestamps or time indices.
[0072] When audio data enters subsequent processing, a time-dimensional sliding segmentation operation needs to be performed on the continuous audio waveform. This operation involves segmenting the audio signal along the audio time axis using continuously moving windows. Each window covers a certain time range; windows can overlap or be adjacent without overlap. The window length determines the time range contained in a single audio segment, and the window movement step determines the time interval between adjacent audio segments. The result of this sliding segmentation is an audio data frame sequence. Each frame in the sequence corresponds to a local time interval on the audio time axis, and multiple local time intervals are arranged sequentially to form a complete sequence. The reason for using time-dimensional sliding segmentation is that sound activity, sound boundaries, and short-term energy changes are all dynamic features within a local time range. Directly processing the entire continuous waveform as a whole would make it difficult to obtain accurate time localization results. The frame-level organization after sliding segmentation provides a unified temporal granularity for subsequent sound boundary localization and short-term energy extraction.
[0073] After the audio data frame sequence is formed, the speech boundary localization operation needs to be performed by an endpoint detection model. The endpoint detection model is used to identify the start and end positions of speech activity in the audio data frame sequence. The input of the endpoint detection model is the audio data frame sequence, and the output is the speech state judgment result or boundary probability result corresponding to each audio frame. The model structure can include a feature input layer, a temporal modeling layer, and a boundary output layer. The feature input layer is used to receive waveform features, spectral features, or energy features of the audio frames; the temporal modeling layer is used to analyze the continuous change relationship between adjacent audio frames, and can use temporal convolutional layers, recurrent layers, or attention layers; the boundary output layer is used to output the judgment result of each audio frame belonging to a speech interval, a silence interval, or a boundary position. If spectral features are used as input, the audio frames can be converted into spectrograms or Mel-frequency features first, and then fed into the convolutional layer to extract local frequency distribution information; if a temporal convolutional structure is used, the convolutional kernel extracts the change pattern between consecutive frames along the time direction; if a recurrent structure is used, the hidden state is used to preserve the speech continuity relationship between consecutive frames. During model training, labeled vocalization interval sample data can be used. The loss function can be classification error loss, and training parameters can include batch size, learning rate, time window length, and hidden dimension. After the endpoint detection model completes inference, the system locates the start and end positions of vocalization in the audio data frame sequence based on the output results. Then, it merges consecutive audio frames in a vocalization state to obtain valid sound segments. Valid sound segments represent continuous intervals with clear speech activity on the audio timeline. These intervals can subsequently serve as important bases for video-side time alignment and label generation.
[0074] For example, the expression for the set of valid sound segments is:
[0075] in, Indicates the start time of the i-th valid sound segment; This indicates the end time of the i-th valid sound segment; i represents the sequence number of the valid sound segment. {} represents a single valid sound segment; {} represents a set consisting of all valid sound segments.
[0076] In fintech business environments, endpoint detection models can be deployed in video auditing systems for remote account opening, remote face-to-face verification, dual-recording verification, or transaction confirmation. The model input is a sequence of audio data frames from the customer's voice interaction, and the output is the judgment result of the vocalization state of each audio frame, thereby extracting the valid sound segments corresponding to the customer's answers to questions or voice confirmations. In healthcare business environments, endpoint detection models can be deployed in remote psychological counseling, remote rehabilitation guidance, or online health communication platforms. The model input is a sequence of audio data frames from the client, and the output is the judgment result of the vocalization state of each audio frame, used to locate the client's vocal expression range. The input data format remains consistent in both scenarios—a time-ordered audio frame sequence—and the output data format also remains consistent—boundary judgment results used to generate valid sound segments.
[0077] After determining the effective sound segments, it is necessary to extract the short-time energy amplitude of each frame of audio data from the audio data frame sequence. The short-time energy amplitude represents the acoustic energy intensity of a single audio frame within its local time range. The calculation of short-time energy is based on the amplitude of the audio sampling points. Each frame of audio data contains several consecutive sampling points. The system performs energy accumulation processing on the amplitude values of these sampling points, that is, squares the amplitude and sums it within the frame, and then normalizes it according to the frame length to obtain the energy amplitude of the audio data of that frame. The short-time energy amplitude can reflect the change in sound intensity within the current frame. Audio frames with obvious sound production usually have higher energy amplitudes, while audio frames that are close to silence or have weak sound production usually have lower energy amplitudes. In subsequent processing, the short-time energy amplitude is not only used to assist in identifying vocal activity, but also to characterize the trend of energy change from weak to strong and from strong to weak during speech expression.
[0078] After extracting the short-time energy amplitude, it needs to be normalized. The purpose of normalization is to eliminate the influence of different acquisition devices, different environmental noise levels, and different speech volumes on the energy value scale, allowing energy values between different audio segments to be compared within a uniform numerical range. Normalization can be performed using maximum value normalization, dividing the short-time energy amplitude of each frame by the maximum energy value in the current audio segment; or using mean-variance normalization, subtracting the mean and dividing by the standard deviation to normalize the energy distribution to a stable range. The normalized short-time energy amplitudes retain their temporal order. Multiple normalized energy values are organized sequentially according to the arrangement order of the audio data frame sequence, forming an audio frame-level energy sequence. Each energy value in the audio frame-level energy sequence corresponds to an audio frame time position, and this sequence reflects the continuous process of speech intensity changing over time. The audio frame-level energy sequence works in conjunction with the effective sound segments. The effective sound segments are used to determine the speech activity intervals, while the audio frame-level energy sequence is used to describe the intensity changes within and at the edges of the intervals, providing more granular temporal information for subsequent video time mapping and soft tag construction.
[0079] This embodiment achieves simultaneous acquisition of speech activity intervals and frame-level energy change information by performing modal separation, sliding segmentation, speech boundary localization, and short-time energy extraction on the multimedia data stream. Valid sound segments are used to determine the time range of speech presence, while audio frame-level energy sequences are used to describe changes in speech intensity, enabling audio information to have a more detailed expressive ability in the time dimension, thereby providing a more accurate temporal basis for subsequent video side tag generation.
[0080] In one embodiment, step S20 above includes: S201, perform a first duration forward extension operation on the starting endpoint of the effective sound segment, and perform a second duration backward extension operation on the ending endpoint of the effective sound segment to generate a target effective segment. The extension range of the first duration forward extension operation is greater than the extension range of the second duration backward extension operation. S202, perform image parsing operation on the video data to obtain a video frame image sequence, perform positioning operation on the video frame image sequence, and extract the target region contained in the video data; S203, obtain the preset basic motion probability within the time range corresponding to the target effective segment, perform a weighted merging operation on the preset basic motion probability and the audio frame-level energy sequence to obtain a gradient probability value sequence, assign the gradient probability value sequence to the time node corresponding to the target region in the video frame image sequence, and generate the initial soft label sequence of the target region contained in the video data.
[0081] In this embodiment, the effective sound segment represents the phonation time interval already identified on the audio side, and the time interval includes a start endpoint and an end endpoint. The start endpoint corresponds to the time position when the speech activity enters the phonation state, and the end endpoint corresponds to the time position when the speech activity exits the phonation state. Using only the effective sound segment as the basis for subsequent tag generation can easily lead to the shortening of the motion start and end positions on the video side, because the motion changes in the target area do not completely coincide with the changes in the speech signal. The phonation preparation stage usually occurs earlier than the stable phonation stage, and the closure recovery stage usually ends later than the stable phonation stage; therefore, boundary adjustment of the effective sound segment is necessary. The boundary adjustment adopts an asymmetric boundary expansion method, the purpose of which is to form an effective time range on the time axis that is closer to the actual motion changes.
[0082] A first-duration forward extension operation is performed on the starting endpoint of the effective sound segment, representing the expansion of the initial boundary along the time axis before the sound emission occurs. The first duration covers the preparatory movement time of the target area that has already occurred before the sound emission. A second-duration backward extension operation is performed on the ending endpoint of the effective sound segment, representing the expansion of the ending boundary along the time axis after the sound emission ends. The second duration covers the remaining movement time of the target area after the sound emission ends. The extension range of the first-duration forward extension operation is greater than that of the second-duration backward extension operation, indicating that the time expansion is not a symmetrical translation, but is set separately according to the continuous movement characteristics before and after the sound emission. This setting allows for a wider coverage range in the initial stage and a shorter but still effective coverage range in the ending stage. After forward and backward extensions, the original effective sound segment is converted into the target effective segment. The target effective segment is wider in time than the effective sound segment, but it is still distributed around the actual sound emission activity and does not expand without boundaries.
[0083] The first and second durations can be set using preset parameters or based on statistical data from the business scenario. In fintech remote interviews or remote account opening videos, the target subject typically makes a clear preparatory opening movement when answering confirmation questions. The initial extension can be set more fully to ensure that the target area movement before the start of speech enters the label range. The final extension can remain smaller to cover the closing movement after the answer. In medical and health remote consultation videos, the target subject's expression pace may be slower. Both the initial and final extensions can be adjusted based on the frame rate of the acquisition device, speaking speed, and local movement duration. The parameters can be provided using a fixed duration configuration or a time window configuration based on the sampling rate or video frame rate. If an adaptive configuration is used, the first duration can be set based on the average time difference between the start of movement and the start of speech in historical samples, and the second duration can be set based on the average time difference between the end of speech and the return to stillness.
[0084] For example, the formula for asymmetric expansion of the target effective fragment is:
[0085] in, Indicates the start time of the i-th valid sound segment; This indicates the end time of the i-th valid sound segment; Indicates the forward extension duration from the starting endpoint; Indicates the duration of the backward extension from the end point; This represents the expanded target valid fragment; This indicates that the forward extension is greater than the backward extension, reflecting asymmetric expansion.
[0086] Once the target segment is determined, the corresponding image carrier for that time interval needs to be located on the video side. Image parsing is performed on the video data to decode the encoded video data into image content that can be processed frame by frame. Image parsing operations can include video decapsulation, frame decoding, timestamp extraction, frame order verification, and image format conversion. The decoded result is a video frame image sequence. This sequence represents a set of image frames arranged chronologically, with each frame corresponding to a specific time position. Subsequent localization operations are performed based on this video frame image sequence, rather than directly on the original compressed video stream. This ensures that the spatial location of the target region within each frame can be determined frame by frame.
[0087] The purpose of performing localization operations on a video frame image sequence is to extract local regions directly related to the vocalization action from the entire frame image, i.e., the target region contained in the video data. The target region can be set as the mouth region in the facial local region, or as other local regions that can reflect the vocalization action of the target object. The localization operation can be implemented through a keypoint detection structure, which outputs local facial markers, and then the target region is cropped based on the geometric relationship between the markers. The localization operation can also be implemented through a bounding box regression structure, which directly outputs the boundary parameters of the target region, and then the local region is cropped from the image frame based on the boundary parameters. It can also be implemented through a segmentation structure, which outputs a target region mask, and the image content inside the region is preserved and the background is removed based on the mask. If an intelligent recognition structure is used for localization, the input data is a single frame image in the video frame image sequence, and the output data is the region coordinates, keypoint set, or region mask in the corresponding frame. The structure includes at least an image input layer, a local feature extraction layer, and a region output layer. The image input layer receives the single frame image, the local feature extraction layer extracts contour and texture information, and the region output layer outputs the spatial location result. After arranging the localization results of each frame in chronological order, the set of spatial locations of the target region in the entire video frame image sequence can be obtained.
[0088] A temporal mapping relationship needs to be established between the target valid segment and the video frame image sequence. The target valid segment provides time interval information, while the video frame image sequence provides discrete image frames and their corresponding time positions. The temporal mapping process can be completed based on a unified timestamp or based on frame index and frame rate conversion. The mapping result is represented by the range of video frames covered by the target valid segment. Each frame in this range corresponds to a target region image instance, thus forming the set of time positions for subsequent label assignment. Only by completing this temporal mapping can the time information obtained from the audio side accurately correspond to the specific frame position on the video side, avoiding the direct assignment of time labels to non-corresponding image frames.
[0089] After determining the time range corresponding to the target valid segment, it is necessary to obtain the preset baseline motion probability within that time range. The preset baseline motion probability represents the baseline motion probability for the video frame time position within the coverage area of the target valid segment. The preset baseline motion probability is not the final label value, but rather an initial probability expression used to describe the possibility of motion existing in this time interval. The baseline probability can be in the form of a piecewise constant or in the form of a continuous function that changes gradually over time. If a piecewise constant form is used, different positions within the target valid segment can be assigned different baseline probability values, for example, a higher probability in the central interval and a lower probability in the edge interval. If a continuous function form is used, the baseline probability can gradually increase at the beginning edge and gradually decrease at the end edge, making the time expression smoother. The purpose of setting the preset baseline motion probability is to convert the target valid segment into a probabilistic time expression that can be combined with the audio frame-level energy sequence, so that the subsequent labels are not simply "present within the interval, absent outside the interval," but rather form a continuous change within the interval.
[0090] An audio frame-level energy sequence represents a set of speech energy values arranged in time, with each energy value corresponding to a time position in an audio frame. The energy values reflect the intensity variations of speech over time. Weighting and merging a preset base motion probability with the audio frame-level energy sequence aims to transform speech activity interval information and speech intensity variation information into a more granular label representation. The weighted merging operation can take the form of time-position multiplication, linear combination, or weighted mapping combination. If time-position multiplication is used, the base probability provides the temporal coverage framework, and the energy value provides the local modulation intensity; multiplying them results in higher probability values for high-energy positions and lower probability values for low-energy positions. If linear combination is used, the base probability and energy value are multiplied by independent weights and then summed, which can balance the influence of interval-level and energy-level information. If weighted mapping combination is used, the weights can be adjusted according to the internal position of the target effective segment, so that the central position is more influenced by the energy value, and the peripheral position is more influenced by the base probability. The result of the weighted merging is a gradually changing probability numerical sequence. Each value in the gradual probability numerical sequence corresponds to a time position, and the magnitude of the value indicates the probability that the target area will move at that time position.
[0091] In the term "gradual" in a gradual probability numerical sequence, the label value changes continuously over time, rather than abruptly at boundary positions. The actual motion of a target region typically increases gradually from rest to apparent motion and decreases gradually from apparent motion back to rest. If the label representation uses a single binary form, the temporal edge positions and the center position would be assigned the exact same label state, failing to reflect the true gradual changes in local motion. The gradual probability numerical sequence provides temporal coverage through effective target segments and intensity variation constraints through audio frame-level energy sequences, allowing different values to be obtained at different time positions, forming a smoothly transitioning label representation.
[0092] Assigning a sequence of gradient probability values to the time nodes corresponding to the target regions in the video frame image sequence represents mapping the gradient probabilities obtained from the audio side to specific frame positions in the video side over time. The objects assigned here are not abstract target region names, but rather the corresponding instances of the target regions at various time positions in the video frame image sequence. Each time node corresponds to a target region in one frame, and the values in the gradient probability sequence correspond one-to-one with these time nodes. After assignment, each target region at each time position in the video frame has a corresponding probability value. The probability values at multiple time positions are arranged in the order of the video frames, forming the initial soft label sequence for the target regions in the video data. "Initial" in the initial soft label sequence indicates that the label was generated before subsequent correction processing; "soft label" indicates that the label values use continuous probability values rather than rigid binary judgment; and "sequence" indicates that the label values are arranged sequentially in the time dimension. The initial soft label sequence can be used for subsequent visual recognition structure training, enabling the visual model to learn the motion and change patterns of the target regions over continuous time.
[0093] For example, the formula for generating the initial soft label gradient:
[0094] in, The initial soft tag value represents the time node t; The indicator function representing whether time node t is within the valid speech interval can generally be written as:
[0095] This indicates the preset basic motion probability; The value represents the audio frame-level energy at time node t; k represents the energy adjustment coefficient. () represents a nonlinear mapping function, generally written as:
[0096] t represents a time node.
[0097] This embodiment generates an initial soft tag sequence by asymmetricly expanding the effective sound segments and combining it with the audio frame-level energy sequence. This allows the time tags to more closely approximate the actual motion range of the target area. Using different amplitude expansions at the start and end helps cover local motion changes before and after sound emission. The weighted merging of the preset base motion probability and the audio frame-level energy sequence ensures continuous temporal variation of the tag values. Therefore, the initial soft tag sequence can simultaneously reflect both temporal range information and intensity change information, thereby reducing deviations caused by boundary truncation and abrupt tag changes.
[0098] In one embodiment, step S30 above includes: S301, Extract a local image sequence containing the target region; S302, the local image sequence is input into a pre-trained visual extraction network to perform spatial feature extraction operations and generate a frame-level spatial feature vector sequence; S303, input the frame-level spatial feature vector sequence into the temporal network to perform temporal context extraction operation, and generate a visual feature sequence; S304, The visual feature sequence is input into the output network layer to obtain the predicted numerical sequence; S305, determine the cross-entropy loss metric between the predicted numerical sequence and the initial soft label sequence; S306, extract the difference magnitude between the predicted values corresponding to adjacent frames in the predicted value sequence, and establish a smoothing regularization term based on the difference magnitude; S307, Combine the smoothing regularization term with the cross-entropy loss metric to generate the target loss; S308, the network parameters of the pre-trained visual extraction network, the temporal network, and the output network layer are adjusted using the target loss. When the preset training stopping condition is met, the network parameter adjustment is stopped, and the initial motion detection model is obtained.
[0099] In this embodiment, the target region in the video data, after prior localization, corresponds to the same local visual object in consecutive video frames. Extracting a sequence of local images containing the target region means cropping the local image containing the target region from each frame in chronological order, and organizing the local images at each time point into a sequence according to the original chronological order. The purpose of setting up the local image sequence is to compress the input range, reduce interference from background areas, irrelevant textures, and global pose changes on subsequent recognition, and allow the subsequent network to focus on processing local visual information related to the motion state of the target region. Each image in the local image sequence corresponds to a time position; therefore, the sequence preserves the local structure of the target region spatially and the continuous change process of the target region temporally. The cropping process can be completed based on the bounding box parameters, key point position parameters, or region mask parameters of the target region. If a bounding box parameter is used, the cropping module reads the region coordinates in each frame and performs cropping and resizing on the original image. If a keypoint position parameter is used, the cropping module determines the local window range based on the relative geometric relationships between keypoints. If a region mask parameter is used, the cropping module suppresses background areas while preserving the image content within the region. To ensure that local images at different time points can be directly fed into the network, local image sequences typically require resizing, color channel unification, and pixel value normalization. Resizing ensures that all local images have consistent width and height, color channel unification ensures that the input data meets the same network input format, and pixel value normalization reduces numerical fluctuations caused by different devices, lighting conditions, and compression qualities.
[0100] A pre-trained visual extraction network is used to extract a sequence of frame-level spatial feature vectors from a sequence of local images. "Pre-trained" in "pre-trained visual extraction network" means that the network already possesses image feature representation capabilities before entering the current training stage, and can be derived from existing parameter states in image recognition, local region recognition, or video understanding tasks. "Visual extraction" means that the network processes local image content, and the output is a vector representation that expresses image texture, contours, local opening and closing, shadow variations, and edge morphology. The network structure includes at least an input layer, a convolutional feature encoding layer, a normalization layer, a non-linear activation layer, and a feature compression layer. The input layer receives single-frame images from the local image sequence and feeds the image tensor into the convolutional feature encoding layer. The convolutional feature encoding layer extracts edges, lip lines, opening morphology, tooth exposure areas, local shadows, and texture variations from the local image. The normalization layer stabilizes the numerical distribution of features from different batches of images, the non-linear activation layer enhances feature discrimination, and the feature compression layer maps high-dimensional feature maps into fixed-length feature vectors. Since the local image sequence is input frame by frame according to its temporal location, the same network performs spatial feature extraction operations with the same parameters in each frame. Therefore, the output is a sequence of frame-level spatial feature vectors that correspond one-to-one with the temporal location. Each vector in the frame-level spatial feature vector sequence only expresses the spatial image information at that temporal location and does not directly express the temporal change relationship. This sequence provides the input basis for subsequent temporal modeling.
[0101] For example, the expression for a frame-level spatial feature vector sequence is:
[0102] in, This represents the sequence of frame-level spatial feature vectors corresponding to the video. Let T represent a T×D real matrix space; T represents the number of time nodes or video frames; D represents the dimension of the feature vector corresponding to each time node.
[0103] The temporal network inputs a sequence of frame-level spatial feature vectors to perform temporal context extraction, establishing dependencies between consecutive temporal positions and thus expanding the single-frame spatial representation into a temporally related representation. The temporal network processes the sequence of frame-level spatial feature vectors arranged chronologically, outputting a visual feature sequence. Compared to the frame-level spatial feature vector sequence, the visual feature sequence not only retains single-frame image information but also incorporates state change information between preceding and subsequent temporal positions. The temporal network can employ temporal convolutional structures, gated recurrent structures, or temporal attention structures. With a temporal convolutional structure, network layers perform convolution operations on features of adjacent frames along the temporal dimension, extracting opening and closing trends, continuous motion trends, and restoring static trends within a local time window. With a gated recurrent structure, recurrent units record feature changes from historical time positions through a state update mechanism and incorporate these changes into the representation of the current position. With a temporal attention structure, attention weights are used to establish correlations between different time positions, enabling the visual representation of the current position to incorporate context over a larger time range. A temporal network includes at least a sequence input layer, a temporal feature modeling layer, and a temporal output layer. The sequence input layer receives a sequence of frame-level spatial feature vectors, the temporal feature modeling layer completes the temporal dependency representation, and the temporal output layer organizes the results into a visual feature sequence corresponding to the time position. Each time position in the visual feature sequence corresponds to a temporally enhanced feature vector, which reflects both the local visual state of the current image frame and the change pattern within the adjacent time range.
[0104] The visual feature sequence is input into the output network layer to obtain the predicted numerical sequence. This indicates that after spatial feature extraction and temporal context extraction, the high-dimensional temporal representation needs to be further mapped into numerical results related to motion states. The output network layer plays a role in state discrimination and can be composed of fully connected layers, normalization mapping layers, and probability output layers. If the output task is set to continuous probability expression, the output network layer generates a continuous numerical value at each time position, and the magnitude of the value represents the probability of motion in the target region at that time position. If the output task is set to distributed expression, the output network layer generates multiple state components at each time position, forms a state distribution through a normalization function, and then the motion components in the distribution form the predicted numerical sequence. Each predicted value in the predicted numerical sequence is aligned with the corresponding time position in the initial soft label sequence. The predicted numerical sequence is arranged continuously in time and can describe the network's judgment result on the motion process of the local region under the current parameter state. Here, the output network layer, the pre-trained visual extraction network, and the temporal network constitute a complete recognition structure, and the three are connected layer by layer through tensor connections. Local image sequences are processed by a pre-trained visual extraction network to form a frame-level spatial feature vector sequence. This frame-level spatial feature vector sequence is then processed by a temporal network to form a visual feature sequence. Finally, the visual feature sequence is processed by an output network layer to form a predicted numerical sequence. The entire process ensures that the lengths of the input data, feature representations, and output results are consistent in the time dimension.
[0105] The cross-entropy loss metric between the predicted numerical sequence and the initial soft-label sequence is determined, representing the deviation between the model output and the supervision labels that needs to be established during the training process. The initial soft-label sequence is a continuous label representation constructed based on audio temporal information and energy changes, where the label value at each time point represents the soft probability that the target region is in motion. The predicted numerical sequence represents the network's output at the same time point. The cross-entropy loss metric measures the distributional difference between the two sets of values. Here, the cross-entropy metric is not a simple comparison of magnitudes, but rather incorporates the predicted probability and label probability at each time point into a unified loss representation. If the predicted value at a certain time point is close to the label value, then that position contributes less to the total loss; if the difference is large, then that position contributes more to the total loss. Since the predicted numerical sequence and the initial soft-label sequence correspond position-by-position in the time dimension, the cross-entropy loss metric can be used to accumulate the errors at each time point to form the total supervision error. This error is used to guide the model parameters to be updated in a direction that is closer to the label representation.
[0106] For example, the formula for the cross-entropy loss between the predicted numerical sequence and the initial soft-label sequence is:
[0107] in, This represents the cross-entropy loss metric. The initial soft tag value represents the time node t; Represents the predicted value at time point t; log represents the logarithmic operation.
[0108] Extracting the magnitude of the difference between predicted values of adjacent frames in the predicted value sequence and establishing a smoothing regularization term means that during training, not only should the predicted values be close to the initial soft label sequence, but the output changes at adjacent time positions should also maintain reasonable continuity. Target region motion in real videos typically exhibits gradual changes and does not undergo drastic jumps without cause within a very short time. If the training process relies solely on the cross-entropy loss metric, the model may output unstable results at label edges, leading to significant prediction fluctuations in the temporal dimension. By extracting the magnitude of the difference between predicted values of adjacent frames in the predicted value sequence, the output changes between adjacent time positions can be quantitatively characterized. If the difference between predicted values of adjacent frames is too large, the smoothing regularization term is increased; if the change in predicted values of adjacent frames is relatively gradual, the smoothing regularization term is decreased. The smoothing regularization term constrains the temporal continuity of the predicted value sequence, reduces output jitter, and makes the model's predictions smoother at boundary and transition positions. When establishing the smoothing regularization term, the prediction differences at all adjacent time positions can be accumulated, or only adjacent positions within the effective target segment range can be accumulated. Higher weights can also be assigned to time positions near the boundary to enhance the boundary smoothing effect.
[0109] For example, the formula for smoothing the prediction difference between adjacent frames is:
[0110] in, Indicates a smoothing regularization term; This represents the predicted value at time point t. Represents the predicted value at the previous time point; | | represents the magnitude of the difference between predicted values at adjacent time points.
[0111] The target loss is generated by merging the smoothing regularization term and the cross-entropy loss metric. This means that training optimization no longer relies solely on the supervised error, but instead uses a composite loss expression composed of both the supervised error and temporal continuity constraints. In the target loss, the cross-entropy loss metric constrains the predicted results to approximate the initial soft-label sequence, while the smoothing regularization term constrains the output continuity over time. After merging, the target loss simultaneously considers accuracy and stability at different time points. The merging method can use a weighted summation: the cross-entropy loss metric is multiplied by the supervised weight, the smoothing regularization term is multiplied by the smoothing weight, and then summed to obtain the target loss. The supervised weight and the smoothing weight are used to adjust the proportion of the two types of constraints in the total loss. If the smoothing weight is too low, the model is prone to oscillations at boundary positions; if the smoothing weight is too high, the model may excessively pursue smoothness and reduce its sensitivity to real motion changes. The weight values can be set according to the characteristics of business data. For example, in remote interview videos in financial technology, the pace of the answers is usually fast and the time boundary changes are obvious. A medium smoothing weight can be set to balance boundary sensitivity and output stability. In remote consultation videos in medical and health care, the pace of expression may be slower. The smoothing weight can be appropriately increased to enhance the expression of continuous changes.
[0112] For example, the formula for synthesizing the target loss is:
[0113] Indicates target loss; This represents the cross-entropy loss metric. λ represents the smoothing regularization term; λ represents the weight coefficient of the smoothing regularization term.
[0114] The target loss is used to adjust the network parameters of the pre-trained visual extraction network, temporal network, and output network layer. Network parameter adjustment stops when a preset training stopping condition is met, resulting in the initial motion detection model. This indicates that training optimization requires joint updating of the parameters of the complete recognition structure. Although the pre-trained visual extraction network already possesses general image representation capabilities, its parameters still need further adjustment based on the current target region motion recognition task, making spatial features more adaptable to local region motion changes. The temporal network is responsible for extracting temporal relationships, and its parameters determine whether the temporal context information can accurately represent the continuous motion process. The output network layer is responsible for mapping the visual feature sequence to the predicted numerical sequence, and its parameters directly affect the final output result. All three parts of the parameters participate in the backpropagation of the target loss. The parameter adjustment process is typically completed collaboratively by the training control module, the loss calculation module, and the optimization update module. The training control module is responsible for organizing the local image sequence, the initial soft label sequence, and the batch processing samples; the loss calculation module is responsible for calculating the cross-entropy loss metric and the smoothing regularization term based on the predicted numerical sequence, and forming the target loss; the optimization update module performs backpropagation updates on the parameters of the pre-trained visual extraction network, temporal network, and output network layer based on the target loss. Training parameters include at least batch size, learning rate, time window length, feature dimension, smoothing weights, and a training epoch limit. Batch size determines the number of samples used for parameter updates in each epoch; learning rate determines the magnitude of parameter updates; time window length determines the range of consecutive frames in a single input; feature dimension determines the length of the network's internal representation; smoothing weights determine the strength of the temporal continuity constraint; and the training epoch limit restricts the maximum number of training epochs. Preset training stopping conditions can be set based on the number of training epochs, the magnitude of change in target loss, or the output error of validation data. When the target loss decreases below a preset threshold over several consecutive training epochs, or when the training epoch limit is reached, the current parameters are considered to have met the stopping requirements, and the network state at this point is retained as the initial motion detection model. The initial motion detection model maintains structural consistency with the network during training, but its parameters have been specifically optimized for the target region motion recognition task.
[0115] In a fintech business environment, input data can be set as local image sequences and their corresponding initial soft-label sequences from remote face-to-face interviews, remote account openings, credit approvals, and dual-recording quality inspection videos. Output data consists of predicted numerical sequences and an initial motion detection model after training. A pre-trained visual extraction network extracts spatial variation features from local customer images, a temporal network learns continuous motion changes during question-and-answer sessions, and the output network layer outputs time-series motion state predictions. After training with composite loss constraints, the resulting initial motion detection model can be used to identify local motion states in fintech business videos. In a healthcare business environment, input data can be set as local image sequences and their corresponding initial soft-label sequences from remote psychological communication, rehabilitation guidance, and online health exchange videos. Output data is similarly set as predicted numerical sequences and an initial motion detection model after training. The network modules, connections, and parameter update methods remain consistent during model training; only the data source and business usage scenarios of the input samples differ.
[0116] This embodiment extracts a local image sequence of the target region and processes it sequentially through a pre-trained visual extraction network, a temporal network, and an output network layer to obtain a predicted numerical sequence corresponding to the time position. The cross-entropy loss metric constrains the prediction result to be close to the initial soft label sequence, while the smoothing regularization term constrains the output changes at adjacent time positions. The two are combined to form the target loss, and the parameters of the complete network structure are adjusted so that the resulting initial motion detection model has both the ability to recognize motion in the target region and can maintain output stability in the temporal dimension.
[0117] In one embodiment, step S40 above includes: S401, predict the motion state of the visual feature sequence through the initial motion detection model, and extract the motion probability distribution values corresponding to each time node; S402, perform a concatenation and splicing operation on the motion probability distribution values according to the order of occurrence of time nodes to obtain the initial prediction sequence; S403, extract the motion probability distribution values at each time node in the initial prediction sequence, perform information entropy quantization on the motion probability distribution values, and determine the information entropy measurement value corresponding to the motion probability distribution values; S404, Perform a sequence combination operation on the information entropy measurement values according to the order of occurrence of the time nodes to determine the uncertainty sequence of the initial prediction sequence.
[0118] In this embodiment, the initial motion detection model receives the visual feature sequence and outputs the motion state judgment results at each time point. The visual feature sequence represents a set of local visual representations arranged in chronological order, with each visual representation corresponding to a time point in the video. The numerical content comes from the encoding results of the target region image in the spatial and temporal dimensions. When processing this sequence, the initial motion detection model needs to maintain the correspondence between time points. Therefore, the model input typically includes a sequence receiving module and a dimension adaptation module. The sequence receiving module is responsible for reading the visual feature vectors arranged in chronological order, while the dimension adaptation module is responsible for uniformly mapping the input features to the processing dimension within the model, enabling subsequent modules to perform operations in a consistent vector space. If the visual feature sequence comes from a remote face-to-face video in fintech business, the input data can be set as the time-series features of the customer's local area during continuous question-and-answer process; if the visual feature sequence comes from a remote consultation video in healthcare business, the input data can be set as the time-series features of the consultation subject's local area during continuous expression process.
[0119] Predicting motion states from visual feature sequences using an initial motion detection model means the model needs to infer states from the visual representation at each time point without altering the temporal order. Internally, the model can include a temporal modeling module and a distributed output module. The temporal modeling module establishes contextual relationships between consecutive time points, ensuring that judgments at a single time point depend not only on the current visual representation but also on the changing trends at adjacent time points. This module can employ temporal convolutional layers, gated recurrent layers, or temporal attention layers. Temporal convolutional layers extract local change information between adjacent nodes by sliding a convolutional window along the temporal direction, suitable for capturing opening and closing changes and continuous motion changes within a short timeframe. Gated recurrent layers convey information along the temporal direction by hiding states, suitable for expressing continuous motion trends over a longer timeframe. Temporal attention layers assign association weights to different time points, suitable for expressing global dependencies over a large timeframe. The distributed output module maps the output of the temporal modeling module to numerical motion probability distributions. If the target state is simply divided into a static state and a moving state, the distributed output module can output two components at each time point, representing the probability of being static and the probability of being moving, respectively. If the output is designed with multi-level motion intensity division, the distributed output module can output multiple probability components at each time point, corresponding to different motion levels. Regardless of the state division method used, the output result of the distributed output module is a motion probability distribution value that corresponds one-to-one with each time point.
[0120] Extracting the motion probability distribution values for each time node means reading the state distribution at each time node from the model output. This extraction is not a simple copying of values, but rather parsing the distribution results corresponding to each time node from the output tensor one by one, preserving the temporal order information. Each motion probability distribution value at each time node corresponds to a local visual state judgment, and its magnitude reflects the model's tendency towards a motion state at that time node. If the motion probability component at a certain time node is significantly higher than the stationary probability component, it indicates that the time node is closer to a motion state; if the differences between multiple components are small, it indicates that the model's judgment of that time node has strong ambiguity.
[0121] The motion probability distribution values are concatenated in chronological order to obtain the initial prediction sequence. This process organizes the outputs scattered across various time points into a complete temporal prediction representation. The concatenation operation must maintain the original temporal order and cannot alter the node indices or intervals. If a time point is earlier in the original video frame sequence, its corresponding motion probability distribution value must also be earlier in the initial prediction sequence. The concatenated initial prediction sequence thus becomes a sequence data structure consistent with the video timeline. The term "initial" in the initial prediction sequence distinguishes it from subsequent corrected labels or prediction representations; "prediction" indicates that the values originate from model inference; and "sequence" indicates that the results at each time point are arranged sequentially. The initial prediction sequence can be stored as a vector sequence or a matrix, where matrix rows correspond to time points and matrix columns correspond to state components. If the business system only needs a single motion probability output, the motion probability components at each time point can be extracted and arranged sequentially; if the business system needs to retain the complete state distribution, the complete distribution of each time point can be preserved in the initial prediction sequence.
[0122] The motion probability distribution values at each time point in the initial prediction sequence are extracted. This means that after obtaining the complete sequence, the corresponding distribution results are read again on a time-node basis for uncertainty quantification. This process emphasizes that the uncertainty calculation object is still the state distribution at each time point, rather than the overall statistics of the entire sequence. The motion probability distribution value at each time point needs to participate in the quantization process separately because the judgment sharpness at different time points may vary significantly. In continuous video, the motion of the target area usually goes through obvious motion intervals, boundary transition intervals, and stationary intervals. The distribution in obvious motion intervals is often more concentrated, while the distribution in boundary transition intervals is often more dispersed. Therefore, the uncertainty cannot be represented by a uniform constant but needs to be determined one by one on a time-node basis.
[0123] For example, the formula for the information entropy of the initial prediction sequence uncertainty is:
[0124] in, The numerical value representing the information entropy measure at time node t; The numerical value representing the probability distribution of motion at time point t; The log represents the probability value of the complementary state at time point t; log represents the logarithmic operation; t represents the time point.
[0125] Performing information entropy quantization on the probability distribution values indicates the use of information entropy as a measure of uncertainty. Information entropy quantization is performed on the probability distribution at each time point, and the quantization result reflects the dispersion of the model's output at that time point. If the probability distribution at a certain time point is highly concentrated on a single state component, the information entropy value is low, indicating that the model's judgment at that time point is more explicit. If the probabilities of multiple state components at a certain time point are similar, the information entropy value is high, indicating that the model's judgment at that time point is more ambiguous. The input to the information entropy quantization operation is the complete probability distribution at that time point, and the output is a single-valued measurement result. To ensure the comparability of quantization results across different time points, a normalization mapping can be performed after quantization to bring the information entropy value into a uniform numerical range. The closer the normalized information entropy value is to the high value range, the higher the uncertainty at that time point; the closer it is to the low value range, the lower the uncertainty at that time point.
[0126] Determining the information entropy metric corresponding to the motion probability distribution value signifies establishing a fixed correspondence between the quantized information entropy result and the original time nodes. The information entropy metric does not exist independently of the time node but serves as a marker of the stability of the prediction result at that time node. This correspondence has a direct impact on subsequent processing because the subsequent fusion or correction process needs to determine whether to trust the model output based on the degree of uncertainty at a particular time node. If the correspondence is not explicitly saved at the current stage, subsequent stages will be unable to distinguish which time nodes belong to high uncertainty and which belong to low uncertainty. Therefore, after determining the information entropy metric, the system typically needs to record both the time node index and the quantization result.
[0127] The information entropy measurement values are combined in chronological order to determine the uncertainty sequence of the initial prediction sequence. This means integrating the node-by-node information entropy measurement results into a complete temporal uncertainty expression. The sequence combination operation follows the same requirements as the initial prediction sequence assembly, ensuring the time order remains unchanged. Each information entropy measurement value is located at the same time position as its corresponding prediction result, and the length of the combined uncertainty sequence is the same as the length of the initial prediction sequence. This establishes a strict one-to-one correspondence between the initial prediction sequence and the uncertainty sequence. Every prediction result in the initial prediction sequence can be found in the uncertainty sequence with a corresponding uncertainty value. This one-to-one correspondence allows the uncertainty sequence to serve as a direct basis for subsequent label correction or result selection.
[0128] The initial motion detection model is in inference mode during this processing stage. Model parameters remain fixed, and no network parameter adjustments are performed. The focus is on generating state and uncertainty outputs in the time dimension based on the existing model. If a neural network structure is used, the model internally includes at least an input receiving layer, a temporal modeling layer, and a state distribution output layer. The input receiving layer receives the visual feature sequence, the temporal modeling layer is responsible for establishing dependencies between consecutive time nodes, and the state distribution output layer is responsible for outputting the state distribution value at each time node. The input receiving layer and the temporal modeling layer are connected via a feature tensor, and the temporal modeling layer and the state distribution output layer are connected via a time series tensor. During model inference, parameters such as inference batch, time series length, and output distribution dimension can be set. The inference batch determines the number of video segments processed simultaneously, the time series length determines the number of consecutive nodes entering the model at a single time, and the output distribution dimension determines the number of state components at each time node.
[0129] This embodiment uses an initial motion detection model to predict the motion state of a visual feature sequence and further determines an uncertainty sequence. This allows for simultaneous acquisition of motion state judgment results and corresponding degrees of uncertainty on the same timeline. The initial prediction sequence reflects the motion probability distribution at each time point, while the uncertainty sequence reflects the concentration of judgments at each time point. This enables subsequent processing to distinguish between clearly predicted locations and vaguely predicted locations, thereby improving the precision of discrimination in the time dimension.
[0130] In one embodiment, step S50 above includes: S501, Extract the information entropy measurement values distributed at each time node in the uncertainty sequence; S502, Establish node reconstruction sequence; S503, for each time node in the uncertainty sequence, compare the information entropy measurement value corresponding to the current time node with the low-order preset threshold and the high-order preset threshold; S504, if the information entropy measurement value is less than the low-order preset threshold, then extract the prediction value at the current time node in the initial prediction sequence, and assign the prediction value to the position corresponding to the current time node in the node reconstruction sequence. S505, if the information entropy measurement value is greater than the high-order preset threshold, then extract the initial label value of the initial soft label sequence at the current time node, and assign the initial label value to the position corresponding to the current time node in the node reconstruction sequence; S506, if the information entropy measurement value is between the low preset threshold and the high preset threshold, then extract the predicted value at the current time node in the initial prediction sequence and the initial label value at the current time node in the initial soft label sequence, perform a weighted summation operation on the predicted value and the initial label value to generate an intermediate transition value, and assign the intermediate transition value to the position corresponding to the current time node in the node reconstruction sequence. S507, after traversing all time nodes, the reconstructed sequence of the nodes is used as the corrected soft label sequence.
[0131] In this embodiment, the uncertainty sequence represents the set of quantization results corresponding to each time node of the initial prediction sequence. Each information entropy metric in the sequence reflects the concentration of the prediction results at the corresponding time node. A lower information entropy metric indicates a more concentrated state distribution of the initial prediction sequence at that time node, resulting in a clearer visual judgment; a higher information entropy metric indicates a more dispersed state distribution of the initial prediction sequence at that time node, resulting in more significant ambiguity in the visual judgment. The role of this sequence is not to provide motion states independently, but rather to provide a basis for the selection and fusion between the initial prediction sequence and the initial soft label sequence. The initial prediction sequence originates from the output of the visual recognition structure, reflecting the visual motion judgment results of the target area over continuous time; the initial soft label sequence originates from the label expression after mapping audio time intervals to energy information, reflecting the prior motion probability within the sound-related time range. Since they differ in origin and reliability at local time locations, differentiated processing is required for each time node using the uncertainty sequence.
[0132] Information entropy measures distributed across various time points in the uncertainty sequence are extracted. These represent quantitative values read node-by-node from the uncertainty sequence, used to determine the credibility of the current node. The purpose of reading node-by-node is to maintain a one-to-one correspondence between time positions, ensuring that the fusion action at each time point is based on the uncertainty of that current node itself, rather than using the average of the entire sequence or overall statistical values. This allows for different strategies to be applied to stable intervals, boundary intervals, and fuzzy intervals. If the object being processed is a remote face-to-face video interview in fintech, the information entropy measure at each time point corresponds to the degree of identification uncertainty of the customer in a local area of a frame; if the object is a remote communication video in healthcare, the information entropy measure at each time point corresponds to the degree of identification uncertainty of the person being consulted in a local area of a frame. The input data format remains consistent, consisting of a sequence of information entropy measures arranged in chronological order.
[0133] A node reconstruction sequence is established, representing a sequence container reserved for the subsequent fusion results, with the number of time nodes matching the number of nodes. Each position in the node reconstruction sequence corresponds to a time node on the original video timeline. The purpose of establishing this sequence is to ensure that the fusion results at different nodes can be written one by one, and to directly form a time-stable output sequence after all nodes have been processed. The length of the node reconstruction sequence is consistent with the uncertainty sequence, the initial prediction sequence, and the initial soft label sequence, with a one-to-one correspondence between position indices. The sequence container can be implemented using a vector array, a time index table, or a frame-level label buffer structure. The specific implementation does not affect the processing logic; it only requires that the node positions remain unchanged during the writing process.
[0134] For each time point in the uncertainty sequence, the information entropy metric corresponding to the current time point is compared with both the lower and higher preset thresholds. This represents the determination of the confidence interval of the prediction result at each time point. The lower preset threshold defines the high confidence interval, and the higher preset threshold defines the low confidence interval. A transition interval is formed between the lower and higher preset thresholds. This comparison process essentially divides the timeline into three states node by node. If the information entropy metric falls below the lower preset threshold, it indicates that the visual prediction at the current node has high stability. If the information entropy metric falls above the higher preset threshold, it indicates that the visual prediction at the current node has strong uncertainty. If the information entropy metric is between the two, it indicates that the current node is neither suitable for relying entirely on visual prediction nor entirely on the initial soft label, but is more suitable for a fusion expression. The thresholds can be given through statistical methods using training data or set based on the information entropy distribution in the business environment. In the fintech business environment, the pace of responses in remote video interviews is relatively fast, and the boundary intervals are relatively concentrated. Setting a threshold can make the low-entropy intervals more prominent and clear in terms of movement location. In the healthcare business environment, the movement in local areas may be more gradual. Setting a threshold can appropriately expand the range of transition intervals in order to retain more information on smooth changes.
[0135] When the information entropy metric is less than a preset low threshold, the predicted value at the current time node is extracted from the initial prediction sequence and assigned to the position corresponding to the current time node in the node reconstruction sequence, indicating that the visual recognition result is prioritized at that time node. This is because low information entropy means that the model output is concentrated at that node, making the visual judgment result clearer. Introducing too many initial soft labels at this point might weaken the already clear judgment of the visual recognition structure. After the predicted value is written into the node reconstruction sequence, the reconstruction result at that time node directly inherits the visual output. This process is suitable for both clearly moving and stable static regions, and is especially suitable for time locations where the target region's morphological changes are significant and local texture changes are clear.
[0136] For example, the formula for predicting inheritance of low-information-entropy nodes is:
[0137] in, This represents the corrected soft tag value at time node t; This represents the initial predicted value at time point t; The numerical value representing the information entropy measure at time node t; The low-order preset threshold t represents the time node.
[0138] When the information entropy metric exceeds a preset high threshold, the initial label value at the current time node is extracted from the initial soft label sequence. This initial label value is then assigned to the position corresponding to the current time node in the node reconstruction sequence, indicating that the label-side representation is preferentially used at that time node. High information entropy indicates that the visual recognition result lacks clarity at the current node, and there may be situations such as similar distributions, fluctuating judgments, or blurred boundaries. In this case, directly using the predicted values from the initial prediction sequence can easily introduce ambiguous judgments into the reconstruction result. The initial soft label sequence originates from audio time information and energy change information. Although it does not come directly from visual recognition, it can provide stable prior constraints at uncertain nodes, making it suitable as alternative information to be written into the node reconstruction sequence. This processing is particularly effective for time nodes near the start and end boundaries of sound emission, because boundary positions often have small visual changes, blurred images, or partial occlusions, while audio-side labels are more stable in terms of temporal range representation.
[0139] For example, the formula for backtracking high-information-entropy node labels is:
[0140] in, This represents the corrected soft tag value at time node t; The initial soft tag value represents the time node t; The numerical value representing the information entropy measure at time node t; The high-order preset threshold t represents the time node.
[0141] When the information entropy metric falls between the lower and higher preset thresholds, the predicted value at the current time node in the initial prediction sequence and the initial label value at the current time node in the initial soft label sequence are extracted. A weighted summation of the predicted and initial label values generates an intermediate transition value, which is then assigned to the position in the node reconstruction sequence corresponding to the current time node. This indicates continuous fusion within a moderate uncertainty range. The intermediate range does not directly use either side's result because it contains both effective visual information and label-side reference value. The weights in the weighted summation operation are determined based on the relative position of the current node's information entropy metric with the lower and higher preset thresholds. The closer the current node is to the lower preset threshold, the higher the weight of the visual prediction; the closer the current node is to the higher preset threshold, the higher the weight of the initial soft label. This intermediate transition value provides a smooth and continuous label change over time, avoiding sudden switching near threshold boundaries. In a physical sense, the intermediate transition value can be understood as a comprehensive probabilistic expression of the motion state of the target area at the current time point, taking into account both the visual judgment result and the prior information of the label.
[0142] For example, the weighted fusion formula for intermediate information entropy nodes is:
[0143] The weighting function can be further written as:
[0144] in, This represents the corrected soft tag value at time node t; This represents the initial predicted value at time point t; The initial soft tag value represents the time node t; The numerical value representing the information entropy measure at time node t; Indicates the preset threshold value for the lower digit; Indicates the high-order preset threshold; This represents the fusion weight of time node t with respect to the initial predicted value; This represents the fusion weight of time node t on the initial soft label value; t represents the time node.
[0145] After traversing all time points, the reconstructed node sequence is used as the corrected soft label sequence. This indicates that after node-by-node judgment, selection, and fusion, the entire node reconstruction sequence has been reconstructed and can be directly used as the corrected label representation for subsequent applications. Compared to the original soft label sequence, the corrected soft label sequence has absorbed information from high-confidence nodes in the visual output; compared to the original prediction sequence, it has eliminated the direct output of unstable nodes through information entropy constraints, thus providing a more balanced label representation in the temporal dimension. The corrected soft label sequence maintains the same length and order as the video frame timeline, with each time point corresponding to a corrected label value. This sequence is suitable for both continuing to participate in the training and updating of the visual recognition structure and for subsequent temporal analysis.
[0146] From an overall operational perspective, the uncertainty sequence provides a basis for node-by-node uncertainty quantification, the node reconstruction sequence provides a node-by-node writing space, and threshold comparison categorizes each time node into high-confidence visual nodes, high-confidence label nodes, and transitional fusion nodes. Predicted values and initial label values enter the reconstruction process within their respective suitable intervals, with intermediate intervals forming a smooth transition through weighted summation. The resulting corrected soft label sequence does not rely on a single source but employs different strategies for different nodes in the time dimension, thus more fully reflecting the continuous changes in the motion state of the target region.
[0147] In a fintech business environment, input data can be set as an initial prediction sequence, an initial soft label sequence, and an uncertainty sequence corresponding to a local area video sequence of the customer. Output data can be set as a corrected soft label sequence for each time node. If the customer's local area motion is clear during a certain question-and-answer period, the information entropy metric of the corresponding node is low, and the corrected soft label sequence retains more of the visual prediction result. If the local area changes are blurry at the beginning or end of the customer's answer, the information entropy metric of the corresponding node may be high or in the middle range, and the corrected result will adopt more of the initial label value or the fusion result.
[0148] This embodiment performs node-by-node fusion processing on the initial prediction sequence and the initial soft label sequence based on the uncertainty sequence. This allows the visual judgment results to be retained at high-confidence nodes, while label-side constraints are introduced at high-uncertainty nodes, forming a continuous transition expression in the intermediate intervals. The resulting corrected soft label sequence balances the discriminative power of the prediction results with the stability of the label information, helping to reduce label bias at boundary and ambiguous positions.
[0149] In one embodiment, step S60 above includes: S601, extract the lower limit true value and the corresponding first time position of the corrected soft tag sequence that are within the preset static value range, and extract the upper limit predicted value and the corresponding second time position of the initial prediction sequence that are greater than the preset motion threshold. S602, compare the first time position with the second time position, and extract video data frames whose time positions overlap to generate difficult sample frames; S603, obtain the initial loss weight corresponding to the difficult example sample frame, and increase the initial loss weight according to a preset increase to generate the difficult example loss weight; S604, Using the corrected soft label sequence as supervision data, and combining it with a loss metric model that includes the hard example loss weights, perform network parameter adjustment operations on the initial motion detection model to obtain a transitional motion detection model; S605, replace the initial motion detection model with the transitional motion detection model, and repeatedly perform motion state prediction, uncertainty sequence determination, fusion processing and network parameter adjustment operations for iterative training. When the preset termination training condition is met, training is stopped to obtain the target motion detection model.
[0150] In this embodiment, the corrected soft-label sequence represents the temporal supervision data after the previous round of label correction. Each label value in the sequence corresponds to a video time position, and the value reflects the probability that the target area at that time position is in motion. The initial prediction sequence represents the prediction results output by the initial motion detection model on the same video time axis under the current parameter state. Each predicted value in the sequence shares the same time position as the corresponding label value in the corrected soft-label sequence. Iterative training of the initial motion detection model does not apply the same training intensity to all sample positions. Instead, it requires identifying difficult-to-distinguish positions from the differences between the corrected soft-label sequence and the initial prediction sequence, so that the model pays more attention to these positions during parameter updates. The preset static value interval is used to describe the range of labels in the corrected soft-label sequence that are close to static, and the preset motion threshold is used to describe the numerical boundaries in the initial prediction sequence that have been judged by the model as being in obvious motion. Extracting the lower limit true value and the corresponding first time position in the corrected soft-label sequence within the preset static value interval represents selecting the time nodes that the label side considers to be close to static from the corrected soft-label sequence. Extract the upper limit predicted values above the preset motion threshold from the initial prediction sequence and their corresponding second time positions. This indicates the locations that the model has strongly judged to be in motion, filtered from the prediction side. The former represents that the label constraint tends to be static, while the latter represents that the visual judgment tends to be in motion. When the two overlap in time position, the node regions that the model is prone to misjudging can be identified.
[0151] The comparison between the first and second time positions transforms numerical conflicts into temporal location results. Only when temporal positions overlap can it be explained that a contradictory phenomenon exists within the same video frame or consecutive time segment: the label side is nearly stationary while the prediction side appears to be moving. Extracting video frames with overlapping temporal positions to generate hard example frames means the system re-labels the original video frames corresponding to these conflicting nodes as high-interest training samples. The significance of hard example frames lies not in simply expanding the training data scale, but in selecting the most valuable parts for correcting model parameters from the existing training data. These samples are typically located in time intervals such as the start and end of vocalization, partial occlusion, small lip opening / closing amplitude, unstable brightness changes, or blurred local motion boundaries. If a uniform loss weight is used directly for training, the model is prone to maintaining its original bias at these sample positions; therefore, these frames need to be weighted separately in subsequent loss calculations.
[0152] For example, the formula for determining difficult sample frames is:
[0153] If both conditions are met at the same time point, the video data frame corresponding to that time point is determined to be a hard case sample frame, that is:
[0154] Wherein, represents the corrected soft tag value at time node t; Indicates a preset static value range; This represents the initial predicted value at time point t; Indicates the preset motion threshold; This represents the video data frame corresponding to time node t; ∧ represents the set of difficult sample frames; ∧ represents the simultaneous satisfaction of two conditions.
[0155] The initial loss weights represent the default weights of the training loss on ordinary sample frames. The initial loss weights corresponding to hard example frames are increased to form hard example loss weights. This adjustment process involves increasing the proportion of error contribution from hard example frames in the loss function, making parameter updates more inclined to correct prediction biases at these locations. The preset increase can be a fixed increment or a proportional amplification. The fixed increment is suitable for environments with stable data scale and relatively uniform sample distribution; the proportional amplification is suitable for environments with large fluctuations in the number of hard example samples and significant differences in local scenes. If the proportion of hard example frames in the entire video is too low, the increase can be moderately increased to provide sufficient constraints on blurred boundary locations during training; if the proportion of hard example frames is too high, it is necessary to avoid excessive weight amplification that could lead to overfitting of the model to a few local samples. After the hard example loss weights are formed, they no longer exist as individual values but are incorporated into the loss measurement model, participating in the error aggregation of each training batch.
[0156] For example, the weighted loss formula for hard example frames is: Let the initial loss weights for ordinary samples be... The loss weight for difficult cases can then be written as:
[0157] Or it can be written in scaled-up form:
[0158] In the weighted loss model, it can be further written as:
[0159] in
[0160] in, Indicates the initial loss weight; Indicates the preset increase range; Indicates the weight of the difficult example loss; Indicates the scaling factor; This represents the total loss including the loss weights for difficult cases; The loss weight represents the time node t; This represents the single-point loss value at time point t. This represents the video data frame corresponding to time node t; This represents the set of difficult example sample frames.
[0161] The loss metric model represents a computational structure that uniformly quantifies the deviation between the prediction results and the supervised data during training. This structure includes at least a supervised error component, a temporal continuity component, and a sample weight component. The supervised error component measures the deviation between the model output and the corrected soft-label sequence, and can be expressed as cross-entropy or squared difference. The temporal continuity component constrains the magnitude of prediction changes between adjacent time points, reducing oscillations in the output at boundary positions after training. The sample weight component applies the hard-example loss weights to the corresponding time points, ensuring that the error of hard-example frames accounts for a higher proportion of the total loss. Thus, the resulting loss metric model is not merely a standard supervised loss, but a training constraint expression with position-differentiated weights. Using the corrected soft-label sequence as supervised data, combined with the loss metric model including hard-example loss weights, the initial motion detection model undergoes network parameter adjustment. This means that during model updates, the corrected soft-label sequence is used as the target output, the hard-example loss weights are used to reinforce high-value samples, and the total loss is used as the basis for parameter updates.
[0162] The initial motion detection model in this training phase includes at least a visual input module, a spatial feature extraction module, a temporal modeling module, and a state output module. The visual input module receives the visual feature sequence of the target region and passes the input tensor to subsequent modules. The spatial feature extraction module further enhances the edge, texture, and morphological change information in the local region. If a relatively stable spatial feature representation has been formed in the previous stage, this module can maintain a small update amplitude here; if there is a significant difference between the current business data and the previous training data, this module needs to participate in joint updates. The temporal modeling module analyzes the dependencies between time nodes, extracting the temporal change representations between the preparatory actions before vocalization, the continuous motion phase, and the recovery to stillness phase. The temporal modeling module can use temporal convolutional layers, gated recurrent layers, or temporal attention layers. The state output module maps the temporal representation to the motion state output per time node. Network parameter adjustment is accomplished through backpropagation, with the goal of making the total loss change in a decreasing direction. During training, the learning rate, batch size, time window length, weight decay coefficient, and hard example weight increment parameters need to be set. The learning rate determines the magnitude of parameter adjustments in each round, the batch size determines the number of samples participating in a single parameter update, the time window length determines the range of consecutive frames entering the model at one time, the weight decay coefficient is used to limit excessive parameter increases, and the hard example weight increment parameter is used to control the proportion of hard example frames in the loss. If the learning rate is too high, the model is prone to training oscillations at hard example frames; if the learning rate is too low, the model's correction speed for hard example regions is insufficient. If the batch size is too small, sample fluctuations will amplify the impact of individual video segments; if the batch size is too large, it will weaken the local correction effect brought by hard example frames.
[0163] After this round of network parameter adjustment, a transitional motion detection model is obtained. The transitional motion detection model maintains the same structure as the initial motion detection model, but its parameter values have absorbed the supervision information from the corrected soft label sequence and hard example frames. Replacing the initial motion detection model with the transitional model means that subsequent training will no longer use the old parameter state, but will instead use the updated model as the starting point for the next round of training. After the replacement, motion state prediction, uncertainty sequence determination, fusion processing, and network parameter adjustment are re-executed, forming a complete iterative update process. This repeated execution is not simply repeating the same set of data processing, but rather that the model regenerates prediction results after each round of parameter updates, recalculates the uncertainty at each time point, re-fuses the uncertainty with the soft labels, and then reconstructs hard example frames and weighted losses based on the new corrected labels and new prediction results. As the rounds progress, the model output gradually approaches the corrected label expression, the number of nodes with high uncertainty gradually decreases, and the distribution of hard example frames also changes. This iterative approach gives the training process self-correction capabilities, rather than remaining in a single round of static training.
[0164] Preset training termination conditions are used to control when to stop iterative training. Termination conditions can be based on an upper limit on the number of training epochs, or on indicators such as the magnitude of change in total loss, the magnitude of change in validation error, or the percentage change in the number of hard example frames. If the decrease in total loss is below a preset threshold after several consecutive training epochs, it indicates that the current parameter state has stabilized; if the recognition error on the validation set no longer decreases, it indicates that the benefits of continued training have significantly diminished; if the number of hard example frames no longer decreases significantly after several consecutive epochs, it indicates that the model's correction of high-conflict locations is nearing saturation. Training stops after the preset training termination conditions are met, the current parameter state is retained, and the target motion detection model is obtained. Compared to the initial motion detection model, the target motion detection model has stronger adaptability at hard example locations, boundary locations, and low-amplitude motion locations because these locations receive additional weight constraints and repeated corrections during iterative training.
[0165] In a fintech business environment, input data can be set as visual feature sequences of target areas from videos used for remote face-to-face interviews, remote account openings, credit approvals, and dual-recording verifications. Supervisory data consists of corrected soft-label sequences, auxiliary data is the initial prediction sequence, and loss-weighted data is the hard-example loss weights. The model output is the motion state judgment results at continuous time points. Since customer answers in remote face-to-face interview scenarios often have clear question-and-answer boundaries, hard-example sample frames are more likely to be concentrated near the beginning and end of the answer; therefore, iterative training has strong value in correcting boundary positions. In a healthcare business environment, input data can be set as visual feature sequences of target areas from videos used for remote consultations, rehabilitation follow-ups, or online psychological communication, with supervisory and auxiliary data maintaining the same format. Because consultation videos have a slower pace, hard-example sample frames may appear continuously in long periods of weak motion; therefore, the training parameters can be appropriately adjusted by changing the time window length and the hard-example weight increase to allow the model to more fully learn slow motion changes.
[0166] This embodiment identifies difficult sample frames based on the corrected soft-label sequence and the initial prediction sequence, and increases the training weights of these sample frames in the loss metric model. This allows iterative training to more effectively correct prediction biases at high-conflict time locations. After the transitional motion detection model replaces the old model, it repeatedly performs prediction, quantization, fusion, and parameter adjustment, which can continuously reduce errors at uncertain locations, making the resulting target motion detection model more stable in its representation of motion states at boundary and ambiguous locations.
[0167] In one embodiment, a target region motion state detection device is provided, which corresponds one-to-one with the target region motion state detection method described in the above embodiments. (Refer to...) Figure 3 , Figure 3This is a schematic diagram of the functional modules of a preferred embodiment of the target area motion state detection device of the present invention. The modules include: audio / video preprocessing module 10, soft tag generation module 20, initial training module 30, uncertainty evaluation module 40, tag fusion correction module 50, iterative optimization module 60, and inference detection module 70. Detailed descriptions of each functional module are as follows: The audio and video preprocessing module 10 is used to acquire audio data and video data synchronized with the audio data, and to extract effective sound segments and audio frame-level energy sequences from the audio data; The soft tag generation module 20 is used to perform asymmetric boundary expansion on the effective sound segment to obtain the target effective segment, and generate an initial soft tag sequence for the target region contained in the video data based on the target effective segment and the audio frame-level energy sequence. The initial training module 30 is used to extract the visual feature sequence of the target region in the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain an initial motion detection model. Uncertainty assessment module 40 is used to predict the motion state of the visual feature sequence through the initial motion detection model to obtain an initial prediction sequence and determine the uncertainty sequence of the initial prediction sequence. The label fusion correction module 50 is used to fuse the initial prediction sequence and the initial soft label sequence according to the uncertainty sequence to obtain the corrected soft label sequence; The iterative optimization module 60 is used to iteratively train the initial motion detection model based on the corrected soft label sequence and the initial prediction sequence. When the preset termination training condition is met, the training is stopped to obtain the target motion detection model. The inference detection module 70 is used to acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0168] Specific limitations regarding the target area motion state detection device can be found in the aforementioned limitations regarding the target area motion state detection method, and will not be repeated here. Each module in the aforementioned target area motion state detection device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0169] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a target area motion state detection method on the server side.
[0170] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements the functions or steps of a target area motion state detection method on the client side.
[0171] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data; The effective sound segment is asymmetrically extended to obtain the target effective segment, and an initial soft tag sequence of the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence. Extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain the initial motion detection model; An initial prediction sequence is obtained by predicting the motion state of the visual feature sequence using the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined. The initial prediction sequence and the initial soft label sequence are fused based on the uncertainty sequence to obtain the corrected soft label sequence. The initial motion detection model is iteratively trained based on the corrected soft label sequence and the initial prediction sequence. Training stops when a preset termination condition is met, and the target motion detection model is obtained. Acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0172] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data; The effective sound segment is asymmetrically extended to obtain the target effective segment, and an initial soft tag sequence of the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence. Extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain the initial motion detection model; An initial prediction sequence is obtained by predicting the motion state of the visual feature sequence using the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined. The initial prediction sequence and the initial soft label sequence are fused based on the uncertainty sequence to obtain the corrected soft label sequence. The initial motion detection model is iteratively trained based on the corrected soft label sequence and the initial prediction sequence. Training stops when a preset termination condition is met, and the target motion detection model is obtained. Acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
[0173] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0174] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0175] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0176] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0177] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for detecting the motion state of a target region, characterized in that, Includes the following steps: Acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data; The effective sound segment is asymmetrically extended to obtain the target effective segment, and an initial soft tag sequence of the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence. Extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain the initial motion detection model; An initial prediction sequence is obtained by predicting the motion state of the visual feature sequence using the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined. The initial prediction sequence and the initial soft label sequence are fused based on the uncertainty sequence to obtain the corrected soft label sequence. The initial motion detection model is iteratively trained based on the corrected soft label sequence and the initial prediction sequence. Training stops when a preset termination condition is met, and the target motion detection model is obtained. Acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
2. The target region motion state detection method as described in claim 1, characterized in that, Acquire audio data and video data synchronized with the audio data, and extract effective sound segments and audio frame-level energy sequences from the audio data, including: Receive a multimedia data stream, perform a modality separation operation on the multimedia data stream, and obtain audio data and video data synchronized with the audio data; Perform a time-dimension sliding segmentation operation on the audio data to obtain an audio data frame sequence; The audio data frame sequence is subjected to a sound boundary localization operation using an endpoint detection model to generate valid sound segments. Extract the short-time energy amplitude of each frame of audio data in the audio data frame sequence, and normalize the short-time energy amplitude to obtain the audio frame-level energy sequence of the audio data.
3. The target area motion state detection method as described in claim 1, characterized in that, The effective sound segment is asymmetrically extended to obtain the target effective segment, and an initial soft label sequence for the target region contained in the video data is generated based on the target effective segment and the audio frame-level energy sequence, including: A first duration forward extension operation is performed on the starting endpoint of the effective sound segment, and a second duration backward extension operation is performed on the ending endpoint of the effective sound segment to generate a target effective segment. The extension range of the first duration forward extension operation is greater than the extension range of the second duration backward extension operation. An image parsing operation is performed on the video data to obtain a video frame image sequence, and a localization operation is performed on the video frame image sequence to extract the target region contained in the video data. Obtain the preset basic motion probability within the time range corresponding to the target effective segment, perform a weighted merging operation on the preset basic motion probability and the audio frame-level energy sequence to obtain a gradient probability value sequence, assign the gradient probability value sequence to the time node corresponding to the target region in the video frame image sequence, and generate the initial soft label sequence of the target region contained in the video data.
4. The target area motion state detection method as described in claim 1, characterized in that, Extract the visual feature sequence of the target region from the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain an initial motion detection model, including: Extract a local image sequence containing the target region; The local image sequence is input into a pre-trained visual extraction network to perform spatial feature extraction operations, generating a frame-level spatial feature vector sequence. The frame-level spatial feature vector sequence is input into a temporal network to perform temporal context extraction operations, generating a visual feature sequence. The visual feature sequence is input into the output network layer to obtain the predicted numerical sequence; Determine the cross-entropy loss metric between the predicted numerical sequence and the initial soft label sequence; Extract the magnitude of the difference between the predicted values of adjacent frames in the predicted value sequence, and establish a smoothing regularization term based on the magnitude of the difference; The smoothing regularization term is combined with the cross-entropy loss metric to generate the target loss. The target loss is used to adjust the network parameters of the pre-trained visual extraction network, the temporal network, and the output network layer. When the preset training stopping condition is met, the network parameter adjustment is stopped, and the initial motion detection model is obtained.
5. The target area motion state detection method as described in claim 1, characterized in that, An initial prediction sequence is obtained by predicting the motion state of the visual feature sequence using the initial motion detection model, and the uncertainty sequence of the initial prediction sequence is determined, including: The motion state of the visual feature sequence is predicted by the initial motion detection model, and the motion probability distribution values corresponding to each time node are extracted. The motion probability distribution values are concatenated and spliced in chronological order of occurrence to obtain the initial prediction sequence. Extract the motion probability distribution values at each time node in the initial prediction sequence, perform information entropy quantization on the motion probability distribution values, and determine the information entropy measurement values corresponding to the motion probability distribution values. The information entropy measurement values are combined in chronological order according to the time nodes to determine the uncertainty sequence of the initial prediction sequence.
6. The target region motion state detection method as described in claim 1, characterized in that, The initial prediction sequence and the initial soft label sequence are fused based on the uncertainty sequence to obtain the corrected soft label sequence, including: Extract the information entropy measurement values distributed at each time node in the uncertainty sequence; Establish a node reconstruction sequence; For each time node in the uncertainty sequence, the information entropy measurement value corresponding to the current time node is compared with the low-order preset threshold and the high-order preset threshold; If the information entropy metric is less than a low-order preset threshold, then the predicted value at the current time node in the initial prediction sequence is extracted, and the predicted value is assigned to the position corresponding to the current time node in the node reconstruction sequence. If the information entropy measurement value is greater than the high-order preset threshold, then the initial label value at the current time node in the initial soft label sequence is extracted, and the initial label value is assigned to the position corresponding to the current time node in the node reconstruction sequence; If the information entropy measurement value is between the low preset threshold and the high preset threshold, then the predicted value at the current time node in the initial prediction sequence and the initial label value at the current time node in the initial soft label sequence are extracted. The predicted value and the initial label value are weighted and summed to generate an intermediate transition value. The intermediate transition value is assigned to the position in the node reconstruction sequence corresponding to the current time node. After traversing all time points, the reconstructed sequence of the nodes is used as the corrected soft label sequence.
7. The target area motion state detection method as described in claim 1, characterized in that, The initial motion detection model is iteratively trained based on the corrected soft-label sequence and the initial prediction sequence. Training stops when a preset termination condition is met, resulting in a target motion detection model, including: Extract the lower limit true value and the corresponding first time position of the corrected soft tag sequence that are within the preset static value range, and extract the upper limit predicted value and the corresponding second time position of the initial prediction sequence that are greater than the preset motion threshold. By comparing the first time position and the second time position, video data frames whose time positions overlap are extracted to generate difficult example sample frames; Obtain the initial loss weight corresponding to the difficult example sample frame, and increase the initial loss weight by a preset increment to generate the difficult example loss weight; Using the corrected soft label sequence as supervision data, and combining it with a loss metric model that includes the hard example loss weights, the network parameters of the initial motion detection model are adjusted to obtain the transitional motion detection model. The initial motion detection model is replaced by the transitional motion detection model. The motion state prediction, uncertainty sequence determination, fusion processing and network parameter adjustment operations are repeatedly performed for iterative training. Training stops when the preset termination condition is met, and the target motion detection model is obtained.
8. A target area motion state detection device, characterized in that, The target area motion state detection device includes: The audio and video preprocessing module is used to acquire audio data and video data synchronized with the audio data, and to extract effective sound segments and audio frame-level energy sequences from the audio data. The soft tag generation module is used to perform asymmetric boundary expansion on the effective sound segment to obtain the target effective segment, and generate an initial soft tag sequence for the target region contained in the video data based on the target effective segment and the audio frame-level energy sequence. An initial training module is used to extract the visual feature sequence of the target region in the video data, and use the visual feature sequence and the initial soft label sequence to train the model to obtain an initial motion detection model. An uncertainty assessment module is used to predict the motion state of the visual feature sequence using the initial motion detection model to obtain an initial prediction sequence and determine the uncertainty sequence of the initial prediction sequence. The label fusion correction module is used to fuse the initial prediction sequence and the initial soft label sequence according to the uncertainty sequence to obtain the corrected soft label sequence; An iterative optimization module is used to iteratively train the initial motion detection model based on the corrected soft label sequence and the initial prediction sequence. Training is stopped when a preset termination training condition is met, and the target motion detection model is obtained. The inference detection module is used to acquire video data to be detected, perform a positioning operation on the video data to be detected to extract the target region to be detected, extract the visual feature sequence to be detected of the target region to be detected, input the visual feature sequence to be detected into the target motion detection model, and obtain the motion state detection result of the target region to be detected.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a target area motion state detection program stored in the memory and executable on the processor. When the target area motion state detection program is executed by the processor, it implements the steps of the target area motion state detection method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a target area motion state detection program, which, when executed by a processor, implements the steps of the target area motion state detection method as described in any one of claims 1-7.