Video content classification and risk early warning method and system based on multi-modal features

Through video content classification and risk warning methods based on multimodal features, the features in video, audio and text data are extracted, and the spatiotemporal overlap and consistency scores are calculated, which solves the problem of difficulty in identifying high-risk behaviors in the existing technology, and achieves efficient and accurate risk warning and video analysis.

CN120071228AActive Publication Date: 2025-05-30BEIJING SMART VISION TECH DEV CO LTD

Patent Information

Application Number
CN202510556908.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-05-30
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Existing video content analysis methods are difficult to effectively identify high-risk behaviors in occlusion, fuzzy or unclear semantic scenarios, and multimodal data processing lacks spatial and temporal consistency and deep mining of behaviors and speech logic, resulting in high false alarm rate of risk identification and insufficient system intelligence level.

Method used

The video content classification and risk warning method based on multimodal features are adopted. By obtaining video frame data, audio data and text data, the position and velocity data of the moving target, the azimuth and pitch angle data of the audio data are extracted, and semantic analysis is carried out to calculate the spatiotemporal overlap between the conflicting behavior segment and the target audio segment, the behavior-sound source consistency score is generated, and the risk determination score and video classification results are finally generated.

Benefits of technology

It realizes multimodal fusion identification and intelligent early warning of potential risk behaviors in video content, significantly improving the accuracy and real-time nature of video analysis, reducing false alarms and missed reports, improving the accuracy of risk judgment, and improving the system's intelligence level in public safety and security monitoring scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071228A_ABST
    Figure CN120071228A_ABST
Patent Text Reader

Abstract

The invention provides a video content classification and risk early warning method and system based on multi-modal features, and relates to the technical field of video processing, and the method comprises the steps: extracting the position and speed data of a moving target in a video frame, generating a motion track diagram and acceleration data, recognizing a conflict behavior segment, and generating the time-space coordinates of the conflict behavior segment; further extracting azimuth angle and pitch angle data in the audio to construct a sound source space distribution diagram, extracting a target audio clip and obtaining sound source position data, and meanwhile, performing semantic analysis on text data to obtain a text feature score; and generating a behavior-sound source consistency score in combination with the space-time coincidence degree of the conflict fragment and the audio fragment, and fusing the behavior-sound source consistency score with a text feature score to calculate a risk judgment score, thereby finally realizing video classification and automatic early warning of a high-risk event.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to video processing technologies, and in particular to a method and system for video content classification and risk warning based on multi-modal features. Background Art

[0002] Currently, with the rapid growth of video surveillance, intelligent security, and Internet video content, a large number of abnormal behaviors and risk events hidden in video data urgently need to be automatically identified and responded to. Traditional video content analysis methods are mostly based on image visual features, and often cannot effectively identify high-risk behaviors in scenes with occlusion, blur, or unclear semantics. Both the accuracy rate and the response speed are difficult to meet the actual combat requirements.

[0003] In recent years, multi-modal learning technologies have gradually been applied to video analysis tasks. By integrating multi-source information such as video images, audio, and text, it helps to comprehensively depict the characteristics of complex events and improve the model's recognition ability for high-risk factors such as conflicting behaviors and dangerous languages. However, most existing methods only process single-modal data, lacking in-depth mining of cross-modal features such as spatio-temporal consistency and the logical association between behaviors and voices, resulting in a high false alarm rate for risk recognition and insufficient system intelligence level.

[0004] Especially in scenarios such as public security monitoring, campus security, and content review, there are a large number of dynamic events such as short-term conflicts and intense disputes, and traditional rules are difficult to cover all potential risks. How to achieve deep integration of multi-modal data, efficient correlation analysis of conflicting behaviors and sound source information, and assist in judging the risk level through semantic information is the key to improving the intelligence level of video warning systems. Therefore, there is an urgent need for a method for video content classification and risk warning based on multi-modal features to achieve accurate recognition and timely response of high-risk events. Summary of the Invention

[0005] Embodiments of the present invention provide a method and system for video content classification and risk warning based on multi-modal features, which can solve the problems in the prior art.

[0006] In the first aspect of the embodiments of the present invention, A method for video content classification and risk warning based on multi-modal features is provided, including: Obtaining video data to be classified and decoding it to obtain video frame data, audio data, and text data; Extracting the position data and speed data of moving objects in the video frame data, generating a motion trajectory map based on the position data, calculating the motion acceleration data based on the speed data, combining the motion trajectory map and the motion acceleration data to identify conflict behavior segments, and generating spatio-temporal coordinate data of the conflict behavior segments; Extract the azimuth data and elevation data from the audio data, construct a sound source spatial distribution map, extract the target audio segment from the sound source spatial distribution map, generate the sound source position data of the target audio segment, and at the same time perform semantic analysis on the text data to generate a text feature score; Calculate the spatio-temporal overlap degree score between the conflict behavior segment and the target audio segment based on the spatio-temporal coordinate data and the sound source position data, and generate a behavior-sound source consistency score; Generate a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score. When the risk determination score exceeds the preset risk threshold, generate warning data and send it to the preset remote monitoring terminal to achieve timely warning response to risk events.

[0007] In an alternative embodiment, Obtain the video data to be classified and decode it to obtain video frame data, audio data and text data, including: Obtain the video data to be classified, extract the video container information in the video data to be classified, parse the global timestamp information from the video container information, and establish an index relationship between the global timestamp information and the video data to be classified; Segment the video data to be classified based on the index relationship into time windows to generate a time window sequence, and perform overlapping processing on the boundary data of adjacent time windows in the time window sequence to generate an overlapping time window sequence; Construct a parallel decoding pipeline, the parallel decoding pipeline includes a video decoder, an audio decoder and a text decoder, input the overlapping time window sequence into the parallel decoding pipeline, decode the video frame data through the video decoder, decode the audio data through the audio decoder, and decode the text data through the text decoder.

[0008] In an alternative embodiment, Extract the position data and speed data of the moving target in the video frame data, generate a motion trajectory map based on the position data, calculate the motion acceleration data based on the speed data, and identify the conflict behavior segment by combining the motion trajectory map and the motion acceleration data to generate the spatio-temporal coordinate data of the conflict behavior segment, including: Perform object detection on the video frame data to obtain the bounding box coordinates, calculate the center point coordinates of the bounding box to obtain the position data of the moving target, and perform filtering processing on the position data to obtain the smoothed position data; Calculate the displacement vector between adjacent frames based on the smoothed position data, and perform modulus operation on the displacement vector to obtain the speed data of the moving target; Generate a motion trajectory graph using the smoothed position data, calculate the curvature features of the motion trajectory graph, extract the coordinates of the trajectory turning points, and at the same time perform a differential operation on the velocity data in the time dimension to obtain motion acceleration data. Screen the motion acceleration data based on the acceleration threshold to determine the acceleration mutation time period; Match the coordinates of the trajectory turning points with the acceleration mutation time period to generate candidate conflict behavior segments, calculate the trajectory continuity coefficient and acceleration change coefficient within the segments, score the candidate conflict behavior segments, and screen to obtain the final conflict behavior segments according to the scoring results; Extract the time stamps and spatial coordinate information corresponding to the final conflict behavior segments to generate the spatio-temporal coordinate data of the conflict behaviors.

[0009] In an alternative embodiment, Extract the azimuth data and elevation data from the audio data, construct a sound source spatial distribution map, and extract the target audio segment from the sound source spatial distribution map. The generated sound source position data of the target audio segment includes: Perform wavelet transform on the audio data to obtain multi-scale band coefficients, construct a band energy matrix based on the multi-scale band coefficients, calculate the inter-channel phase difference according to the band energy matrix, perform weighted combination on the inter-channel phase difference to obtain a phase difference spectrum, and extract the azimuth data and elevation data based on the phase difference spectrum; Perform short-time Fourier transform on the audio data to obtain a time-frequency spectrum, calculate the spatial coherence matrix of the time-frequency spectrum and perform eigenvalue decomposition to obtain a sound source direction vector, construct a time-frequency masker based on the sound source direction vector, and enhance the time-frequency spectrum to obtain an enhanced time-frequency spectrum; Map the azimuth data and elevation data to three-dimensional space coordinates, use the kernel density estimation method to construct a sound source spatial distribution map, and weight the sound source spatial distribution map according to the energy distribution of the enhanced time-frequency spectrum to obtain a weighted sound source spatial distribution map; Perform clustering analysis on the weighted sound source spatial distribution map to obtain an energy aggregation region, mark the target sound source interval in the enhanced time-frequency spectrum according to the energy aggregation region, extract the target audio segment corresponding to the target sound source interval, and generate the sound source position data of the target audio segment.

[0010] In an alternative embodiment, Perform semantic analysis on the text data to generate text feature scores, including: Preprocess the text data to obtain normalized text, count the word frequency features and co-occurrence features from the normalized text, construct a dynamic semantic dictionary, segment the normalized text based on the dynamic semantic dictionary to obtain a word sequence, calculate the semantic association degree between each word in the word sequence and a preset event type to generate word weights, and weight the word sequence according to the word weights to obtain word-level features; Extract scene identification information from the normalized text, calculate the scene dependence degree of the word-level features based on the scene identification information, and weight and adjust the word-level features according to the scene dependence degree to generate sentence-level features; Construct a semantic association graph between sentences, where the nodes are sentence-level features and the edge weights are the semantic similarity between sentences, and perform feature propagation on the semantic association graph to obtain global semantic features; Fuse the global semantic features with the scene dependence degree to generate a document representation vector containing scene information, calculate the similarity between the document representation vector and a preset event vector, and generate a text feature score.

[0011] In an alternative embodiment, Calculating the spatio-temporal overlap degree score between a conflict behavior segment and a target audio segment based on spatio-temporal coordinate data and sound source position data to generate a behavior-sound source consistency score includes: Obtain the correspondence between the behavior position and the sound source position in the historical data, calculate the propagation delay time between the behavior position and the sound source position, and perform time sequence correction on the spatio-temporal coordinate data and the sound source position data according to the propagation delay time; Statistically analyze the behavior-sound source distance distribution and angle distribution in the historical data, construct a spatial weight matrix, and perform spatial correction on the spatio-temporal coordinate data and the sound source position data after time sequence correction by using the spatial weight matrix; Extract the behavior trajectory of the conflict behavior segment and the sound source trajectory of the target audio segment from the spatio-temporal coordinate data and the sound source position data after spatial correction, calculate the spatial overlap degree of the behavior trajectory and the sound source trajectory, and generate a trajectory overlap degree score; Segment the conflict behavior segment and the target audio segment in time, calculate the spatial distance between the behavior position and the sound source position in each time segment, and obtain the corresponding scene occlusion information, calculate the spatio-temporal matching degree of the time segment based on the spatial distance and the scene occlusion information, and generate a segment matching degree score; Weight and fuse the trajectory overlap degree score and the segment matching degree score to generate a behavior-sound source consistency score.

[0012] In an alternative embodiment, Generating a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score includes: Temporally sample the behavior-sound source consistency score and the text feature score to obtain a feature time series, perform a difference operation on the feature time series to obtain a change rate series, detect the mutation moment and the stable moment based on the change rate series, and segment the feature time series according to the peak amplitude at the mutation moment and the duration at the stable moment to obtain multiple feature segments; Extract the time-domain parameters and frequency-domain parameters of the feature segments, allocate the feature segments to different scale levels to generate a feature pyramid; construct a hierarchical transfer matrix using the cross-correlation coefficients between adjacent levels of the feature pyramid to generate a multi-scale feature series; Obtain the area change rate and displacement change rate of the target region in the scene video frame, map them to the corresponding levels of the multi-scale feature series according to the hierarchical transfer matrix to generate scene-adaptive features; Perform temporal statistics and spatial statistics on the scene-adaptive features, and extract the fluctuation coefficient of the temporal statistics and the aggregation coefficient of the spatial statistics; calculate the fluctuation amplitude and fluctuation frequency of the fluctuation coefficient, and generate a risk determination score according to the change trend of the fluctuation amplitude and the density of the fluctuation frequency; Perform spatial grid division on the aggregation coefficient, calculate the feature distribution density and spatial change direction in each grid, determine the behavior occurrence area according to the feature distribution density, determine the behavior development trend according to the spatial change direction, and generate a video classification result based on the behavior occurrence area and the behavior development trend.

[0013] In the second aspect of the embodiments of the present invention, A video content classification and risk warning system based on multi-modal features is provided, including: A first unit for obtaining the video data to be classified and decoding it to obtain video frame data, audio data, and text data; A second unit for extracting the position data and speed data of the moving target in the video frame data, generating a motion trajectory map based on the position data, calculating the motion acceleration data based on the speed data, combining the motion trajectory map and the motion acceleration data to identify conflict behavior segments, and generating spatio-temporal coordinate data of the conflict behavior segments; A third unit for extracting the azimuth angle data and elevation angle data in the audio data, constructing a sound source spatial distribution map, extracting the target audio segment from the sound source spatial distribution map, generating the sound source position data of the target audio segment, and simultaneously performing semantic analysis on the text data to generate a text feature score; A fourth unit for calculating the spatio-temporal overlap degree score between the conflict behavior segment and the target audio segment based on the spatio-temporal coordinate data and the sound source position data to generate a behavior-sound source consistency score; The fifth unit is used to generate a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score, and generate warning data and send it to a preset remote monitoring terminal when the risk determination score exceeds a preset risk threshold, so as to achieve timely warning response to risk events.

[0014] In the third aspect of the embodiments of the present invention, A kind of electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0015] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0016] In this embodiment, it is possible to realize multi-modal fusion recognition and intelligent warning of potential risk behaviors in video content, and significantly improve the accuracy and real-time performance of video analysis. This method combines the motion trajectory and speed change in the video image, and can effectively identify high-risk behaviors such as conflicts and fights; at the same time, the audio azimuth and elevation angle information is introduced to construct a sound source spatial distribution map, so that a spatial association is formed between sound source localization and behavior recognition, effectively enhancing the credibility of conflict behavior judgment. The semantic analysis of text data further supplements the event context information and improves the model's understanding ability of non-visual factors such as language threats and disputes. By calculating the spatio-temporal coincidence degree between the behavior and the sound source, a behavior-sound source consistency score is generated, which helps to reduce false alarms and missed alarms and improve the accuracy of risk determination. Finally, based on the fusion features, a risk determination score and a video classification result are output, which can automatically trigger a warning message when the risk level exceeds a preset threshold and send it to a remote monitoring terminal, improving the system's active response ability and intelligent level in scenarios such as public safety and security monitoring. Description of the Drawings

[0017] Figure 1 It is a flow chart of the method for classifying video content and risk warning based on multi-modal features in the embodiments of the present invention; Figure 2 It is a heat map for generating spatio-temporal coordinate data and locating conflict behaviors in the embodiments of the present invention; Figure 3 It is a bar chart for comparing weighted fusion algorithms in the embodiments of the present invention; Figure 4 It is a structural diagram of the system for classifying video content and risk warning based on multi-modal features in the embodiments of the present invention. Detailed Embodiments

[0018] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] The technical solutions of the present invention will be described in detail below with specific embodiments. These specific embodiments may be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0020] Figure 1 is a schematic flowchart of a method for video content classification and risk warning based on multi-modal features according to an embodiment of the present invention, as Figure 1 shown, the method includes: Obtain the video data to be classified and decode it to obtain video frame data, audio data and text data; Extract the position data and speed data of moving objects in the video frame data, generate a motion trajectory map based on the position data, calculate the motion acceleration data based on the speed data, identify conflict behavior segments by combining the motion trajectory map and the motion acceleration data, and generate spatio-temporal coordinate data of the conflict behavior segments; Extract the azimuth data and elevation angle data in the audio data, construct a sound source spatial distribution map, extract the target audio segment from the sound source spatial distribution map, generate the sound source position data of the target audio segment, and at the same time perform semantic analysis on the text data to generate a text feature score; Calculate the spatio-temporal overlap score between the conflict behavior segment and the target audio segment based on the spatio-temporal coordinate data and the sound source position data, and generate a behavior-sound source consistency score; Generate a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score. When the risk determination score exceeds a preset risk threshold, generate warning data and send it to a preset remote monitoring terminal to achieve timely warning response to risk events.

[0021] In an alternative embodiment, Obtaining the video data to be classified and decoding it to obtain video frame data, audio data and text data includes: Obtain the video data to be classified, extract the video container information in the video data to be classified, parse the global timestamp information from the video container information, and establish an index relationship between the global timestamp information and the video data to be classified; Segment the video data to be classified based on the index relationship into time windows to generate a time window sequence, and perform overlapping processing on the boundary data of adjacent time windows in the time window sequence to generate an overlapping time window sequence; Construct a parallel decoding pipeline, where the parallel decoding pipeline includes a video decoder, an audio decoder, and a text decoder. Input the overlapping time window sequence into the parallel decoding pipeline, and obtain video frame data through decoding by the video decoder, obtain audio data through decoding by the audio decoder, and obtain text data through decoding by the text decoder.

[0022] Exemplarily, when processing the video data to be classified, it is first necessary to obtain the complete video file. The video file usually contains video container format information, which records key information such as video codec format, audio codec format, and subtitle format. The extraction of video container information uses container parsing technology to sequentially read the video file header information block and parse out the video container type identifier. Taking a common MP4 format video as an example, by reading the ftyp box information in the file header, it can be confirmed that the video is encapsulated in the MP4 container format.

[0023] After obtaining the container format information, it is necessary to parse the global timestamp information from the container information. The global timestamp information records the timing relationship of each data frame in the video, including video frame timestamps, audio frame timestamps, and subtitle timestamps. Taking the MP4 format as an example, by parsing the timestamp information table in the mdat box, the display time and decoding time of each data packet can be obtained. These time information are recorded with microsecond-level precision and are used to establish data indexes later.

[0024] To establish the index relationship between the global timestamp and the video data, it is necessary to construct a time index table. This index table contains the mapping relationship from timestamps to data packet positions, and records the specific storage location of the data corresponding to each timestamp in the video file. By traversing the data packets in the video file, record the timestamp of each data packet and its corresponding file offset in the index table. This can achieve fast data positioning and reading based on timestamps.

[0025] After completing the establishment of the index relationship, it is necessary to segment the video data into time windows. The size of the time window is set according to actual application requirements, usually ranging from one second to several seconds. Taking a video frame rate of 30 frames per second as an example, if the time window is set to two seconds, each window will contain 60 video frame data. Divide the video data according to the set time window length to obtain an initial time window sequence.

[0026] To ensure data continuity and avoid information loss, it is necessary to overlap the boundaries of adjacent time windows. The length of the overlapping region is usually set to 20% to 50% of the window length. Taking a two-second time window as an example, a half-second overlapping region is set, that is, adjacent windows share 30 frames of data. This overlapping process can effectively avoid data breaks or information loss at the window boundaries.

[0027] After obtaining the overlapping time window sequence, it is necessary to construct a parallel decoding pipeline. The parallel decoding pipeline consists of three decoders: the video decoder is responsible for processing video stream data, the audio decoder is responsible for processing audio stream data, and the text decoder is responsible for processing subtitles or other text data. The choice of decoder needs to match the encoding format recorded in the video container. Taking a video encoded in H.264 as an example, an H.264 decoder is selected for video frame decoding; for audio encoded in AAC, an AAC decoder is selected for audio data decoding.

[0028] During the parallel decoding process, the data in the overlapping time window sequence is first distributed to the corresponding decoders according to the type. The video decoder receives the compressed video frame data and performs decompression to obtain the original video frames. Taking H.264 encoding as an example, the decoding process includes entropy decoding, inverse quantization, inverse transformation, etc. steps, and finally obtains the video frame data in YUV format. The audio decoder receives the compressed audio data packets and performs decompression to obtain the waveform data. Taking AAC encoding as an example, the decoding process includes bitstream parsing, spectral reconstruction, etc. steps, and finally obtains the audio sampling data in PCM format. The text decoder receives the subtitle data, performs character encoding conversion and format parsing, and obtains the readable text content.

[0029] The decoded video frame data, audio data, and text data maintain a time synchronization relationship, which is convenient for subsequent feature extraction and analysis. The video frame data contains the color and brightness information of the image, the audio data contains the waveform characteristics of the sound, and the text data contains subtitle or annotation information. These data will be used as the basic data input for subsequent behavior analysis and scene understanding.

[0030] In this embodiment, it is possible to achieve efficient parsing, time alignment, and structured preprocessing of video data to be classified, ensuring precise synchronization of multimodal information in the time dimension. By extracting video container format and timestamp information, an accurate data index structure is established, enabling efficient localization and segmented processing of video, audio, and text data according to time windows. At the same time, by setting overlapping time windows and constructing a parallel decoding pipeline, synchronous decoding and unified format processing of video, audio, and text data are achieved, significantly improving data processing efficiency and continuity, and avoiding information loss or behavioral discontinuity at time boundaries. The decoded multimodal data has clear structural features while maintaining time consistency, and can provide accurate and reliable basic data support for subsequent intelligent processing modules such as conflict behavior recognition, sound source analysis, and semantic understanding, effectively improving the accuracy and stability of the entire video risk analysis system.

[0031] In an alternative embodiment, Extract the position data and velocity data of moving objects in the video frame data, generate a motion trajectory map based on the position data, calculate the motion acceleration data based on the velocity data, and identify conflict behavior segments by combining the motion trajectory map and the motion acceleration data. The spatio-temporal coordinate data of the conflict behavior segments includes: Perform object detection on the video frame data to obtain bounding box coordinates, calculate the center point coordinates of the bounding box to obtain the position data of the moving object, and perform filtering processing on the position data to obtain smoothed position data; Calculate the displacement vector between adjacent frames based on the smoothed position data, and perform modulus operation on the displacement vector to obtain the velocity data of the moving object; Generate a motion trajectory map using the smoothed position data, calculate the curvature feature of the motion trajectory map, extract the coordinates of the trajectory turning points, and at the same time perform differential operation on the velocity data in the time dimension to obtain the motion acceleration data. Screen the motion acceleration data based on the acceleration threshold to determine the time period of acceleration mutation; Match the coordinates of the trajectory turning points with the time period of acceleration mutation to generate candidate conflict behavior segments, calculate the trajectory continuity coefficient and acceleration change coefficient within the segment, score the candidate conflict behavior segments, and screen the final conflict behavior segments according to the scoring results; Extract the timestamp and spatial coordinate information corresponding to the final conflict behavior segment to generate the spatio-temporal coordinate data of the conflict behavior.

[0032] Exemplarily, when extracting moving object information from video frame data, it is first necessary to perform object detection on each frame of the image. Object detection uses a deep learning detection model, which can identify the target area in the image and output the bounding box coordinates. The bounding box coordinates include the coordinate values of two points, the upper left corner and the lower right corner, which are used to represent the position and size of the target in the image. Taking the detection of people in a surveillance scenario as an example, the detection model can accurately locate the bounding boxes of each human body area, and can maintain a high detection accuracy even in the case of changing illumination or partial occlusion.

[0033] After obtaining the bounding box coordinates, it is necessary to calculate the center point coordinates of the bounding box. The center point coordinates are obtained by the arithmetic mean of the upper left corner and lower right corner coordinates of the bounding box, representing the centroid position of the target. For a moving target, the sequence of its center point coordinates changes over time to form a motion trajectory. In order to reduce the influence of detection noise on the trajectory, it is necessary to perform filtering processing on the sequence of center point coordinates. The filtering processing uses a moving average method, which performs weighted averaging on the coordinate values of multiple consecutive frames to obtain smoothed position data. Taking a five-frame sliding window as an example, different weights are assigned to the coordinate values within these five frames and summed to obtain the smoothed coordinate value at the current moment.

[0034] Based on the smoothed position data, calculate the displacement vector between adjacent frames. The displacement vector represents the motion direction and distance of the target between two adjacent frames, and is obtained by subtracting the coordinates of the previous frame from the coordinates of the current frame. By calculating the modulus value of the displacement vector, the motion speed of the target at that moment can be obtained. Considering the acquisition frame rate of the video, it is necessary to perform unit conversion on the calculated speed value to obtain the actual motion speed data. For example, in a video with 30 frames per second, the time interval between two frames is 1 / 30 second, and based on this, the displacement can be converted into the actual speed value.

[0035] When generating a motion trajectory graph using the smoothed position data, it is necessary to connect the position points of consecutive frames with a smoothed curve. Calculate the curvature feature of this curve, and the curvature value represents the degree of curvature of the trajectory. Positions with larger curvature values usually correspond to the turning points of the trajectory, and these points are often related to changes in the motion state of the target. By setting a curvature threshold, the key turning point coordinates on the trajectory can be extracted. Taking the example of a pedestrian suddenly changing the motion direction, a large trajectory curvature will be formed at the turning position, and this position is marked as a turning point.

[0036] By performing differential calculation on the speed data in the time dimension, the motion acceleration data can be obtained. The acceleration data reflects the change in the motion speed of the target. By setting an acceleration threshold, the time periods with abnormal acceleration values can be screened out. These time periods usually correspond to the sudden acceleration or sudden deceleration behaviors of the target. For example, when a pedestrian collides or dodges, there will be an obvious speed mutation, resulting in a large acceleration value.

[0037] Perform spatio-temporal matching between the coordinates of the trajectory turning points and the time periods of sudden acceleration changes. If the timestamp of a turning point falls within the time period of sudden acceleration change, then mark the corresponding trajectory segment as a candidate conflict behavior segment. Calculate scoring metrics for each candidate segment, including the trajectory continuity coefficient and the acceleration change coefficient. The trajectory continuity coefficient is obtained by calculating the uniformity of the distances between trajectory points, and a trajectory with good continuity has a higher coefficient value. The acceleration change coefficient is obtained by calculating the fluctuation amplitude of the acceleration values, and a drastic acceleration change results in a higher coefficient value.

[0038] Screen the candidate segments according to the scoring metrics, and determine the segments with both a low trajectory continuity coefficient and a high acceleration change coefficient as the final conflict behavior segments. This screening method can effectively identify abnormal motion behaviors. For example, when two pedestrians have a conflict, their trajectories will show discontinuous turns and significant speed changes at the same time, and such segments will receive a high comprehensive score.

[0039] For the determined conflict behavior segments, extract their corresponding time and space information. The time information includes the start and end timestamps of the segment, and the space information includes the coordinate values of all trajectory points within the segment. This information constitutes the spatio-temporal coordinate data of the conflict behavior and can be used for subsequent behavior analysis and early warning. In this way, the automatic detection and localization of abnormal behaviors in the video are realized, providing important technical support for security monitoring.

[0040] Figure 2 Generate a heat map for the spatio-temporal coordinate data and conflict behavior localization in the embodiments of the present invention, as Figure 2As shown, this figure presents the heatmap analysis results of spatio-temporal coordinate data generation and conflict behavior localization. Within a spatial range of 8m × 6m, the system accurately records and visualizes the spatial distribution of conflict behaviors. In the figure, circular areas of different sizes and transparencies represent the frequency and intensity of conflict behaviors. Among them, the highest frequency of conflict behaviors occurred in Area A (coordinates approximately (2.5m, 1.5m)), reaching 24 times, with the deepest transparency; Area B (coordinates approximately (3.5m, 3.0m)) followed, with 18 conflict occurrences; Areas C and D had 12 and 9 conflict occurrences respectively, belonging to the medium-frequency conflict regions; Area E had only 5 conflict occurrences, belonging to the low-frequency conflict region. Two typical conflict behavior trajectories are also marked in the figure, including the trajectory from coordinates (2.0m, 1.2m) to (2.5m, 1.5m), with time from 2.4 seconds to 4.7 seconds; and the trajectory from coordinates (3.0m, 2.5m) to (3.5m, 3.0m), with time from 8.2 seconds to 11.5 seconds. This demonstrates that the proposed technical solution has significant advantages in the spatio-temporal localization of conflict behaviors, can more accurately capture the occurrence time and spatial location of conflict events, provides strong technical support for security monitoring and early warning, and is particularly suitable for the safety management of crowded places.

[0041] In this embodiment, a deep learning model is used for object detection, which can accurately locate the target position in complex scenarios, and smooth the motion trajectory through moving average filtering, significantly reducing the interference of detection noise on motion analysis. By combining the smoothed position data with the velocity change information, a complete and continuous motion trajectory map can be constructed, and key turning points and velocity mutation moments in the trajectory can be extracted through curvature and acceleration analysis to achieve effective identification of abnormal motion behaviors. By setting double screening indicators of trajectory continuity and acceleration change, conflict behavior segments with high-risk characteristics can be accurately screened out, reducing false positives and false negatives. The finally extracted spatio-temporal coordinate data provides accurate input for subsequent audio-video fusion analysis and early warning decision-making, effectively improving the response speed and intelligent level of the video surveillance system in detecting sudden behaviors and preventing risks.

[0042] In an alternative embodiment, Extract the azimuth data and elevation data from the audio data, construct a sound source spatial distribution map, and extract the target audio segment from the sound source spatial distribution map. The generated sound source position data of the target audio segment includes: Perform wavelet transform on the audio data to obtain multi-scale band coefficients, construct a band energy matrix based on the multi-scale band coefficients, calculate the inter-channel phase difference according to the band energy matrix, weight and combine the inter-channel phase differences to obtain a phase difference spectrum, and extract the azimuth data and elevation data based on the phase difference spectrum; Perform a short-time Fourier transform on the audio data to obtain a time-frequency spectrum, calculate the spatial coherence matrix of the time-frequency spectrum and perform eigenvalue decomposition to obtain a sound source direction vector, construct a time-frequency masker based on the sound source direction vector, and enhance the time-frequency spectrum to obtain an enhanced time-frequency spectrum; Map the azimuth data and elevation angle data to three-dimensional space coordinates, construct a sound source spatial distribution map using the kernel density estimation method, and weight the sound source spatial distribution map according to the energy distribution of the enhanced time-frequency spectrum to obtain a weighted sound source spatial distribution map; Perform clustering analysis on the weighted sound source spatial distribution map to obtain an energy aggregation region, mark the target sound source interval in the enhanced time-frequency spectrum according to the energy aggregation region, extract the target audio segment corresponding to the target sound source interval, and generate the sound source position data of the target audio segment.

[0043] Exemplarily, when processing audio data to obtain sound source spatial information, it is first necessary to perform time-frequency domain analysis on the original audio data. Multi-scale frequency band decomposition can be achieved through wavelet transform, which decomposes the audio signal into sub-band signals in different frequency ranges. The wavelet transform uses orthogonal wavelet basis functions to perform progressive decomposition on the audio data to obtain coefficients reflecting different frequency characteristics. Taking the speech signal collected by an eight-channel microphone array as an example, performing five-layer wavelet decomposition on the audio data of each channel can obtain multiple band coefficients covering from low frequency to high frequency.

[0044] Based on the obtained multi-scale band coefficients, it is necessary to construct a band energy matrix. The band energy matrix records the energy distribution of different frequency bands at each time point. By calculating the sum of squares of each band coefficient, the energy value of the band is obtained. For each time window, calculate the energy values of all frequency bands to form a band energy matrix. For example, when processing a human voice signal in environmental noise, the main energy of the human voice is concentrated in the middle frequency band, and these frequency bands will show higher values in the energy matrix.

[0045] When calculating the inter-channel phase difference according to the band energy matrix, it is necessary to select adjacent microphone pairs for processing. By comparing the signal phases of two channels in the same frequency band and at the same time point, the time difference of the sound wave arriving at different microphones can be obtained. Weight and combine the phase differences of all microphone pairs, and the weight value is proportional to the band energy, so as to highlight the contribution of the frequency bands with stronger energy. The combined phase difference spectrum reflects the position information of the sound source in space.

[0046] When extracting azimuth data and elevation data from the phase difference spectrum, a spatial spectrum estimation method is adopted. This method is based on the law of sound wave propagation and converts the phase difference into spatial angle information. The azimuth represents the position of the sound source in the horizontal plane, and its value range is from zero degrees to three hundred and sixty degrees; the elevation represents the position of the sound source in the vertical plane, and its value range is from negative ninety degrees to positive ninety degrees. For example, when the speaker is directly in front of the microphone array, the azimuth is close to zero degrees and the elevation is close to zero degrees.

[0047] At the same time, perform a short-time Fourier transform on the audio data to obtain a time-frequency spectrum reflecting the time-frequency characteristics of the sound signal. The short-time Fourier transform uses a Hanning window function, the window length is set to twenty-five milliseconds, and the overlap rate is set to fifty percent. This can achieve a good balance between time and frequency resolution. For the spectrum data of each time window, calculate the spatial coherence matrix between different channels. The spatial coherence matrix reflects the correlation between the signals of each channel.

[0048] Perform eigenvalue decomposition on the spatial coherence matrix to obtain the direction vector of the main sound source. Eigenvalue decomposition decomposes the spatial coherence matrix into eigenvalues and eigenvectors, and the eigenvector corresponding to the largest eigenvalue is the direction of the main sound source. Based on this direction vector, construct a time-frequency masker to suppress the noise in the interference direction. The time-frequency masker assigns different weights to each time-frequency point in the time-frequency spectrum, thereby achieving the enhancement of the target sound source.

[0049] Convert the azimuth data and elevation data into point coordinates in a three-dimensional space rectangular coordinate system. When converting, the distance estimation from the sound source to the microphone array needs to be considered. Adopt the kernel density estimation method to construct a three-dimensional probability density distribution centered on these spatial points to obtain the sound source spatial distribution map. The kernel density estimation uses a Gaussian kernel function, and the kernel width parameter is adaptively adjusted according to the sparsity of the spatial points.

[0050] When weighting the sound source spatial distribution map, use the energy distribution of the enhanced time-frequency spectrum as the weight. The higher the energy of the time-frequency spectrum in a region, the higher the weight obtained for the corresponding spatial position. This weighting method can highlight the position information of the main sound source. For example, in a scenario where multiple people are speaking, the speaker with a louder voice will form an obvious energy aggregation region in the weighted spatial distribution map.

[0051] Perform clustering analysis on the weighted sound source spatial distribution map, and use the density clustering algorithm to identify the energy aggregation regions. The density clustering algorithm can discover clusters of any shape and has good robustness to noise points. By setting the density threshold and the minimum cluster size, scattered noise points can be filtered out and the main sound source aggregation regions can be retained.

[0052] According to the identified energy aggregation regions, mark the corresponding time and frequency ranges in the enhanced time-frequency spectrum. These ranges are the target sound source intervals. Extract the audio segments corresponding to these intervals, and at the same time record the spatial position information of the sound source to generate the sound source position data of the target audio segments. For example, for a speaker in a meeting scenario, their speaking segments can be accurately extracted and their specific positions in the meeting room can be marked.

[0053] In this embodiment, by combining wavelet transform and short-time Fourier transform, the time-frequency characteristics of the audio signal are comprehensively characterized, and by fusing the phase difference spectrum and spatial coherence matrix analysis, the azimuth and elevation angle information of the sound source are accurately extracted. Using spatial spectrum estimation and eigenvalue decomposition methods, the target sound source can be effectively enhanced and environmental noise can be suppressed, realizing sound source separation in complex scenarios. Further combining kernel density estimation and density clustering algorithms, the energy aggregation regions in the spatial distribution of the sound source are constructed and identified, so as to extract the target audio segments with clear spatio-temporal labels. This scheme not only improves the accuracy of sound source recognition and localization, but also has strong noise suppression and multi-source separation capabilities, and is applicable to various scenarios such as meeting recording, monitoring analysis, and speech interaction, providing key support for acoustic space perception and speech intelligent processing.

[0054] In an alternative embodiment, Perform semantic analysis on the text data, and the generated text feature scores include: Preprocess the text data to obtain normalized text, count the word frequency features and co-occurrence features from the normalized text, construct a dynamic semantic dictionary, segment the normalized text based on the dynamic semantic dictionary to obtain a word sequence, calculate the semantic association degree between each word in the word sequence and a preset event type to generate word weights, and weight the word sequence according to the word weights to obtain word-level features; Extract the scene identification information from the normalized text, calculate the scene dependence degree of the word-level features based on the scene identification information, and weight and adjust the word-level features according to the scene dependence degree to generate sentence-level features; Construct a semantic association graph between sentences, where the nodes are sentence-level features and the edge weights are the semantic similarities between sentences, and perform feature propagation on the semantic association graph to obtain global semantic features; Fuse the global semantic features with the scene dependence degree to generate a document representation vector containing scene information, calculate the similarity between the document representation vector and a preset event vector, and generate text feature scores.

[0055] Exemplarily, when performing semantic analysis on text data, it is first necessary to preprocess the original text to obtain normalized text. The preprocessing includes multiple steps: removing special characters and invalid characters, uniformly converting all punctuation marks to standard formats, handling homophones and variant forms of Chinese characters, and unifying the character encoding format. For example, for video subtitle text, it is necessary to remove timestamp markers, remove duplicate characters, correct typos, etc. For social media text containing emoticons or special symbols, these symbols need to be converted into corresponding text descriptions or directly deleted.

[0056] After obtaining the normalized text, it is necessary to count the word frequency features and co-occurrence features. The word frequency features reflect the number of times each word appears in the text and its distribution, and the co-occurrence features reflect the mutual relationship between words. For example, in the text describing a fight event, words such as "shove" and "dispute" often appear simultaneously, and this co-occurrence relationship contains important semantic information. Based on the word frequency features and co-occurrence features, a dynamic semantic dictionary is constructed. The dynamic semantic dictionary is different from the traditional static dictionary. It can dynamically adjust the semantic relationship of words according to the current text content and better adapt to specific scenarios.

[0057] Use the constructed dynamic semantic dictionary to segment the normalized text. The segmentation process adopts the maximum matching principle and also considers the context relationship of words. For ambiguous words, disambiguation is performed according to the semantic information of other words in their context. For example, in the sentence "They had a fight event", "fight" is recognized as a whole word instead of being wrongly segmented into the two words "hit" and "frame". The word sequence obtained by segmentation preserves the word order information of the original text.

[0058] For each word in the word sequence, calculate its semantic association degree with the preset event types. The preset event types include multiple typical scenarios, such as fight events, theft events, etc. The calculation of the semantic association degree is based on the word vector space model. Each word is mapped to a vector in a high-dimensional semantic space, and the semantic association degree is obtained by calculating the distance between the word vector and the event type vector. This association degree serves as the weight of the word and is used for subsequent feature extraction.

[0059] Weight the word sequence according to word weights to obtain word-level features. The word-level features retain the original semantic information of the words and highlight the keywords related to the target event through weights. For example, in a text describing a fight event, words such as "fight" and "brawl" will obtain higher weights, while irrelevant words such as "weather" and "time" will obtain lower weights. For example, in a video surveillance scenario, taking a video data containing a fight event as an example, its corresponding text data may come from the subtitle information of the video or the surveillance record. The original text may be like this: "At 3:25 pm, in the community square, two men had a dispute over a parking issue. One of the men was emotional and shouted abuse. Subsequently, the two sides shoved each other and it evolved into a fight. The surrounding people called the police." When preprocessing this text, first remove the timestamp format "3:25 pm" and uniformly convert it to the standard time format. The processed normalized text is more concise and standardized: "In the community square, two men had a dispute over a parking issue. One of the men was emotional and shouted abuse. Subsequently, the two sides shoved each other and it evolved into a fight. The surrounding people called the police." When counting the word frequency and co-occurrence features, it is found that words such as "dispute", "shove", and "fight" often co-occur in the description of fight events. The co-occurrence relationships of these words are recorded in the dynamic semantic dictionary. The word segmentation process cuts the text into a word sequence: "community / square / on / two / men / because / parking / issue / happen / dispute / one / of / them / man / emotion / agitate / loud / shout / abuse / subsequently / both / sides / happen / shove / and / evolve / into / mutual / fight / surrounding / people / call / police / handle". When calculating the semantic relevance degree of the words to the preset fight event, words such as "dispute", "shove", "fight", and "shout abuse" obtain higher weights, while words such as "community", "square", and "parking" obtain lower weights. Such weighting reflects the importance of the words for event judgment.

[0060] Extract scene identification information from the normalized text. The scene identification information includes elements such as time, location, and participants, and these information help to understand the specific environment where the event occurs. For example, in the sentence "A fight occurred on the school playground", "school playground" is an important scene identification. Based on these scene identification information, calculate the scene dependence degree of the word-level features. The scene dependence degree reflects the importance of a certain word in a specific scene.

[0061] Perform weighted adjustment on the word-level features according to the scene dependence degree to generate sentence-level features. This step integrates the word-level features with the scene information to make the features better reflect the overall semantics of the sentence. For example, the weight of the word "fight" in the "school" scene may be higher than that in other scenes because fight events in the school environment may require more attention.

[0062] When constructing the semantic association graph between sentences, the sentence-level features of each sentence are used as nodes in the graph. The semantic similarity between sentences is calculated as the weight of the edge, and the similarity calculation takes into account word order information and sentence structure information. For example, although the two sentences "Two people had an argument" and "Both sides started fighting" use different words, they have similar semantics, and the edge weight between them is relatively high.

[0063] Feature propagation is performed on the semantic association graph. The feature propagation process uses a graph neural network model to enable the features of each node to absorb the information of neighboring nodes through multiple rounds of iteration. In this way, global semantic features considering context relationships can be obtained. For example, it may be difficult to determine the event type by looking at the sentence "They started shoving" alone, but in combination with the context "Both sides had an argument" and "Then they started fighting", it can be more accurately understood that this is a description of a fight event.

[0064] The global semantic features are fused with the scene dependence degree to generate a document representation vector. The fusion process adopts an attention mechanism to re-weight the global semantic features according to the scene information. The resulting document representation vector contains both the semantic information of the text and reflects the scene features. For example, in the description of a fight event in a school scene, the document representation vector will retain more features related to the educational environment.

[0065] The similarity between the document representation vector and the preset event vector is calculated to generate a text feature score. The preset event vector is a standardized description of various typical events. By calculating the vector similarity, the matching degree between the text description and the standard event can be obtained. The higher the similarity, the more likely the text describes that type of event. This feature score will be used for subsequent event determination and analysis.

[0066] In this embodiment, through the multi-level semantic analysis of text data, the event-related information in the text can be accurately extracted, and the event type can be accurately identified in combination with scene features. The design using a dynamic semantic dictionary and scene dependence degree improves the system's understanding ability of similar event descriptions in different scenes. Based on the feature propagation mechanism of the semantic association graph, the logical relationship between sentences is effectively captured, enhancing the depth of understanding of the event development process. Incorporating scene information into the text feature analysis process improves the scene adaptability of feature representation, enabling the system to better understand event descriptions in specific environments. Through layer-by-layer feature extraction from the word level to the sentence level and then to the global semantics, a complete text semantic understanding framework is established, improving the accuracy and reliability of event recognition. The finally generated text feature score can accurately reflect the relevance between the text description and the target event, providing a reliable basis for subsequent event determination and analysis.

[0067] In an alternative embodiment, Calculating the spatio-temporal overlap degree score between the conflict behavior segment and the target audio segment based on spatio-temporal coordinate data and sound source position data, and generating the behavior-sound source consistency score includes: Obtain the correspondence between the behavior position and the sound source position in the historical data, calculate the propagation delay time between the behavior position and the sound source position, and perform temporal alignment on the spatio-temporal coordinate data and the sound source position data according to the propagation delay time; Statistically analyze the behavior-sound source distance distribution and angle distribution in the historical data, construct a spatial weight matrix, and perform spatial calibration on the spatio-temporal coordinate data and the sound source position data after temporal alignment by using the spatial weight matrix; Extract the behavior trajectory of the conflict behavior segment and the sound source trajectory of the target audio segment from the spatio-temporal coordinate data and the sound source position data after spatial calibration, calculate the spatial overlap degree of the behavior trajectory and the sound source trajectory, and generate a trajectory overlap degree score; Perform time segmentation on the conflict behavior segment and the target audio segment, calculate the spatial distance between the behavior position and the sound source position in each time segment, and obtain the corresponding scene occlusion information. Calculate the spatio-temporal matching degree of the time segment based on the spatial distance and the scene occlusion information, and generate a segmented matching degree score; Perform weighted fusion on the trajectory overlap degree score and the segmented matching degree score to generate a behavior-sound source consistency score.

[0068] Exemplarily, when performing behavior-sound source consistency analysis, first, it is necessary to obtain the correspondence between the behavior position and the sound source position from the historical data. The historical data contains a large number of labeled behavior events and their corresponding sound source information, which reflects the spatio-temporal correlation characteristics between behaviors and sounds in a specific scenario. For example, in an indoor scenario, due to the influence of wall reflection on sound propagation, there may be deviations in sound source localization, and this characteristic will be reflected in the historical data.

[0069] To accurately calculate the sound propagation delay time, it is necessary to consider the propagation speed of sound waves in the air. Under standard atmospheric pressure and room temperature conditions, the propagation speed of sound waves is approximately 340 meters per second. According to the spatial distance between the behavior position and the sound source position, the theoretical propagation delay time can be calculated. Taking a fight event as an example, when a fight occurs, the sound generated by limb collisions takes a certain amount of time to propagate to the microphone array, and this delay time needs to be compensated in subsequent analysis.

[0070] According to the calculated propagation delay time, perform temporal alignment on the spatio-temporal coordinate data and the sound source position data. The purpose of temporal alignment is to align the behavior occurrence time with the sound acquisition time and eliminate the time difference caused by sound wave propagation. For example, if the fighting location is 10 meters away from the microphone array, the sound propagation takes about 0.03 seconds, and during calibration, the time axis of the sound source position data needs to be shifted forward by this delay amount.

[0071] Next, count the behavior - sound source distance distribution and angle distribution in the historical data. The distance distribution reflects the sound propagation characteristics generated by different types of behaviors, while the angle distribution reflects the accuracy of sound source localization in different directions. In an indoor environment, due to the existence of multipath effects, the sound source localization in some directions may be more accurate than in others. Based on these statistical characteristics, construct a spatial weight matrix, which describes the credibility weights of each position in the space.

[0072] Use the spatial weight matrix to perform spatial correction on the data after temporal correction. Spatial correction takes into account the influence of scene characteristics on sound source localization and weights and adjusts the sound source coordinates at different positions. For example, in the area near the wall, due to the influence of reflected sound, the weight of the sound source localization result will be appropriately reduced; while in the open area, the weight of the sound source localization result is relatively high.

[0073] After completion of the correction, extract the behavior trajectories of the conflict behavior segments from the spatio - temporal coordinate data. The behavior trajectory describes the movement path of the targets involved in the conflict in space. At the same time, extract the sound source trajectories of the target audio segments from the sound source position data. The sound source trajectory reflects the movement of the sound source in space. When calculating the spatial coincidence degree of these two trajectories, it is necessary to consider the matching degree of the distance and direction between the trajectory points.

[0074] In order to analyze the correspondence between behaviors and sound sources in more detail, it is necessary to perform time segmentation on the conflict behavior segments and target audio segments. The length of the segmentation can be set according to actual needs, and usually a value between 0.1 second and 0.5 second is selected. Within each time segment, calculate the spatial distance between the behavior position and the sound source position. At the same time, obtain the corresponding occlusion information from the scene information, including obstacles such as walls and furniture that may affect sound propagation.

[0075] Calculate the spatio - temporal matching degree of the time segment based on the spatial distance and scene occlusion information. When the behavior position and the sound source position are close and there is no significant occlusion in the middle, the matching degree of this segment is relatively high. On the contrary, if the distance is far or there is obvious occlusion, the matching degree is relatively low. For example, when a fight occurs at a corner, due to the occlusion and reflection of the wall, the accuracy of sound source localization will be reduced, and the matching degree in this case needs to be adjusted appropriately.

[0076] Perform weighted fusion on the trajectory coincidence degree score and the segment matching degree score to obtain the final behavior - sound source consistency score. When performing fusion, it is necessary to consider the influence of scene characteristics on the two scores. In an open scene, the weight of the trajectory coincidence degree score can be appropriately increased; in a complex scene, the weight of the segment matching degree score needs to be increased. For example, in a narrow space such as a corridor, due to the limitation of sound wave propagation, the weight of the segment matching degree score should be higher than the weight of the trajectory coincidence degree score.

[0077] Through the above processing flow, the in-depth fusion analysis of behavior data and sound source data is realized. This method fully considers the characteristics of sound wave propagation and scene constraints, and through multi-level spatio-temporal alignment and matching degree calculation, accurately evaluates the consistency degree between behavior and sound source. This analysis method is particularly suitable for abnormal behavior detection in complex scenarios, and can effectively improve the accuracy and reliability of detection.

[0078] Figure 3 This is the bar chart for comparing the weighted fusion algorithms of the embodiments of the present invention. As Figure 3 shown, this bar chart shows the comparison of behavior-sound source consistency scores of different methods in five test scenarios. The white bars represent the present technical solution (based on dynamic weighted fusion), the gray bars represent the traditional average weighted method (such as linear fusion algorithm), and the dark gray bars represent the traditional maximum value method (such as maximum response fusion algorithm). In the static scenario, the performances of the three methods are relatively close, with the present technical solution being 96.8%, and the traditional average weighted and maximum value methods being 92.5% and 93.8% respectively. As the scene complexity increases, the gap gradually expands. In the dynamic scenario, the score of the present technical solution is 94.2%, which is better than 89.5% and 90.2% of the traditional methods; in the multi-source scenario, the score of the present technical solution is 91.5%, significantly exceeding 85.0% and 83.0% of the traditional methods; in the long-distance scenario, 87.0% of the present technical solution is much higher than 80.0% and 77.5% of the traditional methods; in the complex occlusion scenario, the gap further expands, and the present technical solution still maintains a high score of 84.0%, while the traditional methods are only 72.5% and 69.0% respectively. Overall, the present technical solution is on average 8.7 - 11.2 percentage points higher than the traditional methods in various scenarios, and the advantage is more significant especially in complex environments. This method can more accurately reflect the correlation degree between behavior and sound source by dynamically weighting and fusing the trajectory coincidence degree score and the segment matching degree score.

[0079] The prior art usually determines the association between the two by synchronously collecting video and audio data and statically matching the behavior position and the sound source position, but ignores the factors of sound wave propagation delay and dynamic occlusion in complex scenarios, resulting in a high risk of misjudgment in the correspondence between behavior and sound source. This application introduces the calculation of propagation delay time and timing correction to ensure the precise alignment of behavior data and sound source data in the time dimension, improving the basic accuracy of analysis. On this basis, a spatial weight matrix is constructed using historical data to dynamically model the spatial relationship between behavior and sound source, overcoming the limitation of traditional methods that only rely on the straight-line distance and ignore the behavior-sound source angle and scene distribution. By extracting and comparing the trajectory coincidence degree of conflicting behaviors and target audio, this application further enhances the judgment ability of the spatio-temporal consistency between the two. Aiming at the problem that it is difficult to accurately quantify the relationship between behavior and sound source in occlusion scenarios, a time-segment evaluation mechanism is proposed, and the matching degree is calculated by combining the spatial distance and occlusion information of each segment, realizing a spatio-temporal matching evaluation with more detailed perception ability. Finally, by fusing multi-dimensional scores to generate a behavior-sound source consistency index, the ability to accurately match conflicting behaviors and audio events in complex environments is significantly improved, meeting the application requirements of higher precision and multi-factor determination.

[0080] In an alternative embodiment, Generating a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score includes: Performing time series sampling on the behavior-sound source consistency score and the text feature score to obtain a feature time series, performing a difference operation on the feature time series to obtain a change rate series, detecting mutation moments and stable moments based on the change rate series, and segmenting the feature time series according to the peak amplitude of the mutation moment and the duration of the stable moment to obtain a plurality of feature segments; Extracting the time domain parameters and frequency domain parameters of the feature segments, allocating the feature segments to different scale levels to generate a feature pyramid; constructing a hierarchical transfer matrix using the cross-correlation coefficients between adjacent levels of the feature pyramid to generate a multi-scale feature series; Obtaining the area change rate and displacement change rate of the target area in the scene video frame, mapping them to the corresponding levels of the multi-scale feature series according to the hierarchical transfer matrix to generate scene adaptive features; Performing time series statistics and spatial statistics on the scene adaptive features, and extracting the fluctuation coefficient of time series statistics and the aggregation coefficient of spatial statistics; calculating the fluctuation amplitude and fluctuation frequency of the fluctuation coefficient, and generating a risk determination score according to the change trend of the fluctuation amplitude and the density of the fluctuation frequency; Perform spatial grid division on the clustering coefficient, calculate the feature distribution density and the spatial change direction within each grid, determine the behavior occurrence area according to the feature distribution density, determine the behavior development trend according to the spatial change direction, and generate a video classification result based on the behavior occurrence area and the behavior development trend.

[0081] Exemplarily, in the actual risk determination and video classification process, it is first necessary to perform temporal sampling on the behavior-sound source consistency score and the text feature score. The sampling interval can be set according to the video frame rate. For example, in a video with 30 frames per second, sampling can be performed every ten frames, which not only ensures the continuity of the data but also reduces the computational amount. The sampled feature time series reflects the dynamic process of event development. For example, in a fight event, the feature score will fluctuate with the change of the conflict intensity.

[0082] Perform a difference operation on the feature time series, calculate the change amount between adjacent sampling points, and obtain the change rate series. The change rate series can reflect the change speed and direction of the feature value. By analyzing the change rate series, the mutation moment when the feature value changes significantly and the stable moment when it remains relatively stable can be detected. For example, when a fight event suddenly escalates, the feature score will show an obvious jump, forming a mutation moment; when the participants are in a continuous confrontation, the feature score is relatively stable.

[0083] Segment the feature time series according to the detected mutation moment and stable moment. At the mutation moment, record the peak amplitude, indicating the severity of the feature value change; at the stable moment, record the duration, indicating the time span during which the feature value remains stable. Based on this information, divide the entire sequence into multiple feature segments. For example, a fight event may include multiple feature segments such as a verbal argument stage, a physical conflict stage, and a stage when the situation calms down.

[0084] Extract time-domain parameters and frequency-domain parameters for each feature segment. The time-domain parameters include statistical features such as the mean and variance, which reflect the overall level and fluctuation of the feature value. The frequency-domain parameters are obtained by performing spectral analysis on the feature sequence and reflect the periodic characteristics of the feature value change. These parameters together constitute a complete description of the feature segment.

[0085] Allocate the extracted feature segments to different levels according to the time scale to construct a feature pyramid. The bottom layer contains the finest-grained feature changes, and the upper layer contains the feature changes with a longer time span. For example, the bottom layer may describe an instantaneous pushing action, while the upper layer describes the development trend of the entire fight process. The hierarchical transfer matrix is constructed by calculating the cross-correlation coefficient between adjacent levels, which describes the correlation relationship between features at different time scales.

[0086] Obtain the area change rate and displacement change rate of the target region from the scene video frames. The area change rate reflects the size change of the target in the image and may be related to the approach or departure of the target; the displacement change rate reflects the movement speed of the target and may be related to the intensity of the target's behavior. These change rates are mapped to the corresponding levels of the feature pyramid according to the hierarchical transfer matrix to generate scene-adaptive features.

[0087] Conduct temporal statistics and spatial statistics analysis on the scene-adaptive features. Temporal statistics calculate the fluctuation coefficient of the feature values over time, and the fluctuation coefficient reflects the intensity of the change of the feature values. Spatial statistics calculate the clustering coefficient of the feature values in the spatial distribution, and the clustering coefficient reflects the spatial distribution characteristics of the feature values.

[0088] Further analyze the fluctuation coefficient to obtain the fluctuation amplitude and fluctuation frequency. The fluctuation amplitude represents the magnitude of the change of the feature values, and the fluctuation frequency represents the frequency of the change of the feature values. According to the change trend of the fluctuation amplitude, the development trend of the event can be judged. For example, if the fluctuation amplitude continues to increase, it may indicate that the conflict is escalating. According to the density of the fluctuation frequency, the urgency of the event can be judged. For example, if the fluctuation frequency suddenly increases, it may indicate that a violent conflict is about to occur.

[0089] Combine the analysis results of the fluctuation amplitude and fluctuation frequency to generate a risk determination score. The calculation of the score takes into account the direction and intensity of the change trend, as well as the density and duration of the change frequency. For example, when continuous high-amplitude and high-frequency fluctuations are detected, the system will give a higher risk determination score, indicating that there may be serious security risks.

[0090] Perform spatial grid division on the clustering coefficient, dividing the scene space into several grid units. Calculate the feature distribution density within each grid unit, which reflects the possibility of abnormal behavior occurring in this area. At the same time, calculate the spatial change direction of the feature values, which reflects the propagation trend of abnormal behavior. For example, in a mass incident, the central area of the event can be determined by analyzing the feature distribution density, and the possible diffusion direction of the event can be predicted by analyzing the spatial change direction.

[0091] Based on the behavior occurrence area determined by the feature distribution density and the behavior development trend determined by the spatial change direction, finally generate the video classification result. The classification result not only includes the determination of the event type, but also includes the prediction of the spatial range and development trend of the event. For example, the system may classify a certain video as "local fight event, occurring in the northeast corner of the scene, with a trend of spreading to the central area".

[0092] In the prior art, static feature fusion or single-scale analysis methods are mostly used in video risk identification and classification, which are difficult to accurately capture the dynamic relationship between behavior and sound source and the impact of scene changes on event evolution, resulting in lagging risk assessment and unstable classification results. In this application, by performing temporal sampling and differential analysis on the behavior-sound source consistency score and text feature score, feature mutation and stable moments are effectively identified, and accurate segmentation of key behavior nodes is achieved; a multi-scale feature pyramid is constructed by combining time-domain and frequency-domain parameters, and feature associations between scales are mined through a hierarchical transfer matrix, enhancing the spatio-temporal expression ability of features. On this basis, the area and displacement change rate of the target region in the scene video are introduced, and they are adaptively mapped to the multi-scale feature hierarchy, fully considering the interference and influence of scene dynamic changes on behavior features, thereby constructing more semantically adaptable scene features. Subsequently, the fluctuation coefficient and aggregation coefficient are extracted through temporal statistics and spatial statistics to respectively characterize the dynamic intensity and spatial concentration degree of behavior features, and then a risk determination score with stronger spatio-temporal perception ability is generated. At the same time, the spatial grid division method is used to refine the feature distribution, and the behavior region and its development trend are analyzed by combining density and change direction. Compared with the coarse-grained event label division method in the prior art, this application significantly improves the recognition granularity and classification accuracy, enhancing the accurate judgment ability of potential risks and behavior types in complex behavior scenarios.

[0093] Figure 4 FIG. is a schematic structural diagram of a video content classification and risk warning system based on multi-modal features according to an embodiment of the present invention, as Figure 4 shown, the system includes: A first unit for acquiring video data to be classified and decoding it to obtain video frame data, audio data, and text data; A second unit for extracting position data and velocity data of moving targets in the video frame data, generating a motion trajectory map based on the position data, calculating motion acceleration data based on the velocity data, identifying conflict behavior segments by combining the motion trajectory map and the motion acceleration data, and generating spatio-temporal coordinate data of the conflict behavior segments; A third unit for extracting azimuth data and elevation data in the audio data, constructing a sound source spatial distribution map, extracting target audio segments from the sound source spatial distribution map, generating sound source position data of the target audio segments, and simultaneously performing semantic analysis on the text data to generate text feature scores; A fourth unit for calculating the spatio-temporal overlap score between the conflict behavior segments and the target audio segments based on the spatio-temporal coordinate data and the sound source position data, and generating a behavior-sound source consistency score; The fifth unit is used to generate a risk determination score and a video classification result based on the behavior-sound source consistency score and the text feature score, generate warning data and send it to a preset remote monitoring terminal when the risk determination score exceeds a preset risk threshold, so as to achieve timely warning response to risk events.

[0094] In the third aspect of the embodiments of the present invention, A kind of electronic device is provided, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.

[0095] In the fourth aspect of the embodiments of the present invention, A computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.

[0096] The present invention can be a method, a device, a system and / or a computer program product. The computer program product may include a computer-readable storage medium, on which computer-readable program instructions for executing various aspects of the present invention are loaded.

[0097] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A video content classification and risk warning method based on multimodal features, characterized in that: include: Obtaining the video data to be classified and decoding it to obtain video frame data, audio data and text data; Extracting the position data and speed data of the moving target in the video frame data, generating a motion trajectory diagram based on the position data, calculating motion acceleration data based on the speed data, identifying the conflict behavior segment by combining the motion trajectory diagram and the motion acceleration data, and generating the spatiotemporal coordinate data of the conflict behavior segment; Extract azimuth data and elevation data from audio data, construct a sound source spatial distribution map, extract the target audio segment from the sound source spatial distribution map, generate the sound source position data of the target audio segment, and perform semantic analysis on the text data to generate a text feature score; Calculate the spatiotemporal coincidence score of the conflicting behavior segment and the target audio segment based on the spatiotemporal coordinate data and the sound source position data to generate a behavior-sound source consistency score; Based on the behavior-sound source consistency score and text feature score, a risk determination score and video classification result are generated. When the risk determination score exceeds the preset risk threshold, warning data is generated and sent to the preset remote monitoring terminal to achieve timely warning response to risk events.

2. The method according to claim 1, characterized in that Obtaining the video data to be classified and decoding it to obtain video frame data, audio data and text data includes: Acquire video data to be classified, extract video container information from the video data to be classified, parse global timestamp information from the video container information, and establish an index relationship between the global timestamp information and the video data to be classified; Based on the index relationship, the video data to be classified is segmented into time windows to generate a time window sequence, and the boundary data of adjacent time windows in the time window sequence are overlapped to generate an overlapping time window sequence; A parallel decoding pipeline is constructed, which includes a video decoder, an audio decoder and a text decoder. The overlapping time window sequence is input into the parallel decoding pipeline, and the video decoder decodes it to obtain video frame data, the audio decoder decodes it to obtain audio data, and the text decoder decodes it to obtain text data.

3. The method according to claim 1, characterized in that Extract the position data and speed data of the moving target in the video frame data, generate a motion trajectory map based on the position data, calculate the motion acceleration data based on the speed data, identify the conflict behavior segment by combining the motion trajectory map and the motion acceleration data, and generate the spatiotemporal coordinate data of the conflict behavior segment, including: Performing target detection on the video frame data to obtain bounding box coordinates, calculating the center point coordinates of the bounding box to obtain position data of the moving target, and performing filtering processing on the position data to obtain smoothed position data; Calculate the displacement vector between adjacent frames based on the smoothed position data, and perform a modulus operation on the displacement vector to obtain the speed data of the moving target; Generate a motion trajectory graph using the smoothed position data, calculate the curvature characteristics of the motion trajectory graph, extract the coordinates of the turning points of the trajectory, and perform a differential operation on the velocity data in the time dimension to obtain motion acceleration data. Filter the motion acceleration data based on the acceleration threshold to determine the acceleration mutation time period; Matching the trajectory turning point coordinates with the acceleration mutation time period to generate candidate conflict behavior segments, calculating the trajectory continuity coefficient and acceleration change coefficient within the segment, scoring the candidate conflict behavior segments, and screening the final conflict behavior segments according to the scoring results; The timestamp and spatial coordinate information corresponding to the final conflict behavior fragment are extracted to generate the spatiotemporal coordinate data of the conflict behavior.

4. The method according to claim 1, characterized in that: Extracting azimuth data and elevation data from the audio data, constructing a sound source spatial distribution map, extracting a target audio segment from the sound source spatial distribution map, and generating sound source position data of the target audio segment include: Performing wavelet transform on audio data to obtain multi-scale frequency band coefficients, constructing a frequency band energy matrix based on the multi-scale frequency band coefficients, calculating inter-channel phase differences according to the frequency band energy matrix, weightedly combining the inter-channel phase differences to obtain a phase difference spectrum, and extracting azimuth angle data and elevation angle data based on the phase difference spectrum; Performing short-time Fourier transform on the audio data to obtain a time-frequency spectrum, calculating a spatial coherence matrix of the time-frequency spectrum and performing eigenvalue decomposition to obtain a sound source direction vector, constructing a time-frequency masker based on the sound source direction vector, and enhancing the time-frequency spectrum to obtain an enhanced time-frequency spectrum; Mapping the azimuth angle data and the elevation angle data into three-dimensional space coordinates, constructing a sound source spatial distribution map using a kernel density estimation method, and weighting the sound source spatial distribution map according to the energy distribution of the enhanced time-frequency spectrum to obtain a weighted sound source spatial distribution map; Cluster analysis is performed on the weighted sound source spatial distribution map to obtain energy concentration areas, and the target sound source interval is marked in the enhanced time-frequency spectrum according to the energy concentration areas, and the target audio segment corresponding to the target sound source interval is extracted to generate sound source position data of the target audio segment.

5. The method according to claim 1, characterized in that Perform semantic analysis on text data to generate text feature scores including: Preprocessing the text data to obtain a normalized text, counting word frequency features and co-occurrence features from the normalized text, constructing a dynamic semantic dictionary, segmenting the normalized text based on the dynamic semantic dictionary to obtain a word sequence, calculating the semantic association between each word in the word sequence and a preset event type to generate a word weight, and weighting the word sequence according to the word weight to obtain a word-level feature; Extracting scene identification information from the normalized text, calculating the scene dependency of the word-level feature based on the scene identification information, and weighting and adjusting the word-level feature according to the scene dependency to generate a sentence-level feature; Constructing a semantic association graph between sentences, where nodes are sentence-level features and edge weights are semantic similarities between sentences, and performing feature propagation on the semantic association graph to obtain global semantic features; The global semantic feature is fused with the scene dependency to generate a document representation vector containing scene information, and the similarity between the document representation vector and a preset event vector is calculated to generate a text feature score.

6. The method according to claim 1, characterized in that The spatiotemporal coincidence score of the conflicting behavior segment and the target audio segment is calculated based on the spatiotemporal coordinate data and the sound source position data, and the behavior-sound source consistency score is generated, including: Obtaining the corresponding relationship between the behavior position and the sound source position in the historical data, calculating the propagation delay time between the behavior position and the sound source position, and performing time sequence correction on the spatiotemporal coordinate data and the sound source position data according to the propagation delay time; Counting the behavior-sound source distance distribution and angle distribution in the historical data, constructing a spatial weight matrix, and using the spatial weight matrix to perform spatial correction on the time-series corrected spatiotemporal coordinate data and sound source position data; Extracting the behavior trajectory of the conflicting behavior segment and the sound source trajectory of the target audio segment from the spatially corrected spatiotemporal coordinate data and the sound source position data, calculating the spatial overlap between the behavior trajectory and the sound source trajectory, and generating a trajectory overlap score; The conflicting behavior segment and the target audio segment are time segmented, the spatial distance between the behavior position and the sound source position is calculated in each time segment, and the corresponding scene occlusion information is obtained, and the spatiotemporal matching degree of the time segment is calculated based on the spatial distance and the scene occlusion information to generate a segment matching degree score; The trajectory coincidence score and the segment matching score are weightedly fused to generate a behavior-sound source consistency score.

7. The method according to claim 1, characterized in that The risk assessment scores and video classification results generated based on the behavior-sound source consistency score and text feature score include: Performing time-series sampling on the behavior-sound source consistency score and the text feature score to obtain a feature time series, performing a differential operation on the feature time series to obtain a change rate series, detecting a mutation moment and a stable moment based on the change rate series, and segmenting the feature time series according to the peak amplitude of the mutation moment and the duration of the stable moment to obtain a plurality of feature segments; Extracting time domain parameters and frequency domain parameters of feature segments, assigning the feature segments to different scale levels, and generating a feature pyramid; constructing a hierarchical transfer matrix using the mutual correlation coefficients between adjacent levels of the feature pyramid, and generating a multi-scale feature sequence; Acquire the area change rate and displacement change rate of the target area in the scene video frame, map them to the corresponding level of the multi-scale feature sequence according to the hierarchical transfer matrix, and generate scene adaptive features; Performing time series statistics and space statistics on the scene adaptive features, extracting the fluctuation coefficient of the time series statistics and the clustering coefficient of the space statistics; calculating the fluctuation amplitude and the fluctuation frequency of the fluctuation coefficient, and generating a risk determination score according to the change trend of the fluctuation amplitude and the density of the fluctuation frequency; The clustering coefficient is divided into spatial grids, and the feature distribution density and spatial change direction are calculated in each grid. The behavior occurrence area is determined according to the feature distribution density, and the behavior development trend is determined according to the spatial change direction. The video classification result is generated based on the behavior occurrence area and the behavior development trend.

8. A video content classification and risk warning system based on multimodal features, used to implement the method described in any one of claims 1 to 7, characterized in that: include: The first unit is used to obtain the video data to be classified and decode it to obtain video frame data, audio data and text data; The second unit is used to extract the position data and speed data of the moving target in the video frame data, generate a motion trajectory diagram based on the position data, calculate motion acceleration data based on the speed data, identify the conflict behavior segment by combining the motion trajectory diagram and the motion acceleration data, and generate the spatiotemporal coordinate data of the conflict behavior segment; The third unit is used to extract the azimuth data and the pitch angle data from the audio data, construct a sound source spatial distribution map, extract the target audio segment from the sound source spatial distribution map, generate the sound source position data of the target audio segment, and perform semantic analysis on the text data to generate a text feature score; The fourth unit is used to calculate the spatiotemporal coincidence score of the conflicting behavior segment and the target audio segment based on the spatiotemporal coordinate data and the sound source position data, and generate a behavior-sound source consistency score; The fifth unit is used to generate risk determination scores and video classification results based on the behavior-sound source consistency score and the text feature score. When the risk determination score exceeds the preset risk threshold, warning data is generated and sent to the preset remote monitoring terminal to achieve timely warning response to risk events.

9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.

10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Target behavior recognition and prediction method in complex dynamic environment

    CN108647582A

  • Fall alarm method, device and equipment

    CN115223331A

  • Multi-modal intelligent analysis method and system based on law enforcement supervision

    CN119229337A

  • System for media correlation based on latent evidences of audio

    US20140009682A1

Cited By

  • Intelligent rail-mounted area safety early warning method and device based on edge calculation

    CN120833656A

  • Public security event intelligent early warning communication system based on AI semantic analysis

    CN120911749A

  • Airport video data real-time analysis system

    CN120976826A

  • Road video event rapid detection method based on edge calculation

    CN121121599A

  • Gravitational acceleration measuring method and system based on YOLOv8

    CN122048994A