Anomaly identification methods, devices, equipment, storage media and products
By dividing and clustering the evaluation video data into stages, and combining the characteristics of the sound source nodes and the evaluation content, the granularity of the action is generated for audio anomaly analysis. This solves the problems of low accuracy and high false alarm rate in anomaly identification in the evaluation scenario, and achieves efficient and accurate anomaly detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中移信息技术有限公司
- Filing Date
- 2026-04-30
- Publication Date
- 2026-07-31
AI Technical Summary
Existing anomaly detection technologies based on video streams suffer from low accuracy and high false alarm rates in bidding evaluation scenarios, and lack global feature analysis capabilities, making it difficult to accurately identify abnormal behaviors during the bidding evaluation process.
By dividing the evaluation video data into stages, clustering is performed based on the fluctuation value of the sound source node range and the similarity of the coding structure to select target video streams. Combined with the criticality of the evaluation content and the layout of the sound source nodes, action granularity is generated to perform audio anomaly analysis to identify abnormal behavior.
It improved the accuracy of anomaly identification in the bidding process, reduced the false alarm rate, and enabled detailed analysis of key links and efficient filtering of routine links, thereby improving detection efficiency and accuracy.
Smart Images

Figure CN122493361A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent video surveillance technology, and in particular to an anomaly identification method, apparatus, device, storage medium, and product. Background Technology
[0002] As a core component of bidding activities, the compliance and transparency monitoring of the bid evaluation process is crucial to ensuring fairness and impartiality. With the integration of video surveillance and AI recognition technologies, anomaly detection based on video streams has become an important means of supervising the bid evaluation process. However, the bid evaluation process has significant stage-specific characteristics, with marked differences in the semantic features of the images and the behavioral patterns of personnel at different stages, placing high demands on the global adaptability and scenario customization of anomaly detection.
[0003] However, existing anomaly detection technologies based on video streams suffer from the following main problems: On the one hand, they rely on sequential event streams to analyze local scene features, lacking the ability to comprehensively analyze the global features of the evaluation scene, making it difficult to accurately identify truly abnormal behavior. On the other hand, most of these technologies require manual video segmentation and anomaly labeling before AI inference and recognition, which is not only inefficient but also leads to inconsistent recognition standards due to differences in human experience, further affecting the accuracy of anomaly judgment.
[0004] In summary, existing anomaly detection methods based on video streams suffer from low anomaly recognition accuracy and high false alarm rate. Summary of the Invention
[0005] This invention provides an anomaly identification method, apparatus, device, storage medium, and product to solve the problems of low anomaly identification accuracy and high false alarm rate in existing video stream-based anomaly detection methods.
[0006] According to one aspect of the present invention, an anomaly identification method is provided, comprising: Based on the bidding stage to which the video belongs, the bidding video data from multiple channels is divided to obtain a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream; Based on the fluctuation values of the sound source node range and the similarity of the coding structure of each segment of the video stream, the first video stream set is determined by clustering. Based on the audio quality relationship of each sound source node in each segment of the video stream, the second video stream set is determined by clustering. Determine the target video stream that simultaneously belongs to both the first video stream set and the second video stream set; For any GOP group in the target video stream as the current GOP group, the action granularity of the current GOP group is generated based on the criticality of the evaluation content corresponding to the current GOP group. Based on the sound source node layout of the current GOP group, the action granularity is adaptively adjusted to obtain the target action granularity, and audio anomaly analysis is performed in combination with the spatial audio characteristics of the current GOP group. Based on the results of the audio anomaly analysis and the target action granularity, anomaly identification is performed on the current GOP group.
[0007] According to another aspect of the present invention, an anomaly detection device is provided, comprising: The video segmentation module is used to segment multiple bidding video data according to the bidding stage to which the video belongs, to obtain a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream; The set determination module is used to cluster and determine the first video stream set based on the fluctuation value of the sound source node range and the similarity of the coding structure of each segment video stream, and to cluster and determine the second video stream set based on the audio quality relationship of each sound source node in each segment video stream. The target video stream determination module is used to determine a target video stream that simultaneously belongs to both the first video stream set and the second video stream set. The granularity determination module is used to take any GOP group in the target video stream as the current GOP group and generate the action granularity of the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group. The audio analysis module is used to adapt the action granularity based on the sound source node layout of the current GOP group to obtain the target action granularity, and to perform audio anomaly analysis in combination with the spatial audio characteristics of the current GOP group. An anomaly identification module is used to identify anomalies in the current GOP group based on the results of the audio anomaly analysis and the target action granularity.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the anomaly identification method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the anomaly identification method according to any embodiment of the present invention.
[0010] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the anomaly identification method described in any embodiment of the present invention.
[0011] The technical solution of this invention involves dividing multiple bidding video data streams into segments based on the bidding stage to which the video belongs, resulting in a set of segmented video streams, wherein each segmented video stream set includes at least one segmented video stream. A first video stream set is determined by clustering based on the fluctuation values of the sound source node ranges and the similarity of the coding structure of each segmented video stream. A second video stream set is determined by clustering based on the audio quality relationships of each sound source node in each segmented video stream. A target video stream that simultaneously belongs to both the first and second video stream sets is identified. For any Group of Points (GOPs) in the target video stream, the current GOP is selected as the current GOP. Based on the criticality of the bidding content corresponding to the current GOP, an action granularity for the current GOP is generated. Based on the sound source node layout of the current GOP, the action granularity is adaptively adjusted to obtain a target action granularity. Audio anomaly analysis is then performed in conjunction with the spatial audio features of the current GOP. Anomaly identification is performed on the current GOP based on the results of the audio anomaly analysis and the target action granularity. First, based on the segmented multi-channel evaluation video data from the evaluation stage, a first video stream set is determined by combining the fluctuation values of the sound source node range and the similarity of the coding structure. Simultaneously, a second video stream set is determined by utilizing the audio quality relationships between sound source nodes. The target video stream is located through intersection filtering, and static segments are eliminated to reduce false alarms. Second, for the GOP group of the target video stream, action granularity is generated based on the criticality of the evaluation content, and adaptive adjustments are made according to the stability of the sound source node layout. Audio anomaly analysis is performed in conjunction with spatial audio features. Finally, anomaly identification is performed based on the analysis results and the target action granularity. This achieves synergy between fine analysis of key links and efficient filtering of routine links, thereby improving accuracy and reducing the false alarm rate.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1This is a flowchart of an anomaly identification method provided in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of video stream partitioning provided in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structure of an anomaly identification device provided in Embodiment 2 of the present invention; Figure 4 This is a schematic diagram of the structure of an electronic device provided in Embodiment 3 of the present invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] Example 1 Figure 1 This is a flowchart of an anomaly identification method provided in Embodiment 1 of the present invention. This embodiment is applicable to anomaly detection based on multiple bidding evaluation videos at the bidding evaluation site, determining anomalies in the bidding evaluation. This method can be executed by an anomaly identification device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes: S110. Based on the bidding evaluation stage to which the video belongs, the bidding evaluation video data of multiple channels are divided to obtain a segmented video stream set, wherein the segmented video stream set includes at least one segmented video stream.
[0018] In this embodiment, the bidding evaluation stage to which the video belongs refers to different work stages of the bidding evaluation activity divided according to a preset business process. In some embodiments, the bidding evaluation stage to which the video belongs includes at least one of the following: capital verification stage, bid splitting stage, bid reading stage, review stage, negotiation stage, and comprehensive conclusion stage. The bidding evaluation video data is a video stream collected from the bidding evaluation site and containing both video and audio information. A segmented video stream can be understood as a video segment obtained by dividing the bidding evaluation video data according to the bidding evaluation stage to which the video belongs on the timeline. A set of segmented video streams can be understood as the set of all segmented video streams obtained by dividing the synchronously collected multiple bidding evaluation video data according to the bidding evaluation stage to which the video belongs.
[0019] Specifically, multiple video cameras deployed at the bid evaluation site collect multi-channel bid evaluation video data for comprehensive analysis. Then, based on the bid evaluation stage to which the video belongs, the multiple bid evaluation video data channels are divided into segmented video streams for different bid evaluation stages; these segmented video streams form a set of segmented video streams.
[0020] For example, suppose there are i channels of bid evaluation video data, and the bid evaluation stage to which the video belongs includes three stages: capital verification, bid splitting, and price announcement. Based on the bid evaluation stage to which the video belongs, the multiple bid evaluation video data channels are divided into segments to obtain an initial set of segmented video streams. , can be represented as: ; R1, R2, and R3 represent the capital verification stage, the bid splitting stage, and the price announcement stage, respectively. This represents the evaluation video data for the i-th path.
[0021] Furthermore, based on the sound source node layout of each initial segmented video stream, each initial segmented video stream is marked. The resulting set of segmented video streams after marking can be represented as: ; Among them, the sound source node layout is a spatial distribution layout of personnel at the bidding site based on video image recognition, which can be mainly divided into three typical categories: dual-track form, single-sided form, and closed-loop form. This data records the quantitative ratios of distance and clustering, representing the overall spatial distribution of all bidders in the different layout configurations described above. The configuration refers to the distribution of bidders; a double-track configuration involves two rows of people sitting side-by-side, a single-sided configuration involves a single row of people sitting, and a closed-loop configuration involves a circular arrangement of people sitting.
[0022] Figure 2 This is a schematic diagram of video stream partitioning provided in Embodiment 1 of the present invention. Figure 2As shown, there are three video streams for bid evaluation. The bid evaluation stages to which the videos belong include the capital verification stage, the bid splitting stage and the price announcement stage, the review stage and the negotiation stage. Based on the bid evaluation stage to which the videos belong, the multiple video streams are divided into initial segmented video stream sets. Furthermore, according to the sound source node layout of each initial segmented video stream, including dual-track, single-sided and closed-loop forms, each initial segmented video stream is marked, resulting in a segmented video stream set.
[0023] In some embodiments, the sound source node corresponds to the bid evaluation personnel at the bid evaluation site; the node layout is one of a double-track form, a single-sided form, and a closed-loop form, wherein the double-track form includes a layout in which bid evaluation personnel are arranged in two parallel rows, the single-sided form includes a layout in which bid evaluation personnel are arranged in a single row, and the closed-loop form includes a layout in which bid evaluation personnel are arranged in a ring.
[0024] S120. Based on the fluctuation value of the sound source node range and the similarity of the coding structure of each segmented video stream, cluster to determine the first video stream set. Based on the audio quality relationship of each sound source node in each segmented video stream, cluster to determine the second video stream set.
[0025] In this embodiment, the sound source node represents the location identifier of the evaluator in the evaluation video data. This location identifier has both visual location attributes and spatial sound source location attributes. The sound source node range can be understood as the structured observation area of a single evaluator in the visual image of the evaluation video data. The sound source node range fluctuation value refers to a comprehensive quantitative index obtained by calculating the area change rate of the sound source node range of each sound source node (corresponding to each evaluator) in a single segmented video stream across multiple consecutive frames, and then aggregating the change rates of all nodes. This index is used to efficiently identify and filter candidate segmented video streams where the dynamics of the image are significantly abnormal due to the entry and exit of groups of evaluators or intense individual activities. The coding structure similarity refers to the degree of consistency between the coding structure features of each group of pictures (GOP) within a single segmented video stream. The coding structure features are constructed based on the total jump distribution statistics of each B-frame within a GOP relative to its preceding and following reference frames. The total jump count is the sum of the forward and backward jump counts of each B-frame. Higher coding structure similarity indicates more consistent coding patterns in the video segments, and a lower likelihood of post-editing or malicious interruption. By identifying the regularity of inter-frame reference relationships in video coding, segmented video streams with abnormal inter-frame correlations due to post-editing or transcoding are excluded, ensuring the content of the selected segmented video streams is coherent and free from malicious tampering. Audio quality relationships are the comparison between the multi-dimensional audio features corresponding to each sound source node and global statistics under the same layout, used to quantitatively assess whether there is illegal information transmission or abnormal prompts between sound source nodes. These multi-dimensional audio features include, but are not limited to, loudness, brightness, intensity of specific frequency components, pitch, and speech rate. The first video stream set can be understood as a subset obtained after visual feature filtering of the segmented video stream set. The second video stream set can be understood as a subset obtained after audio feature filtering of the segmented video stream set.
[0026] Specifically, for each segmented video stream in the segmented video stream set, the range of each sound source node in the current segmented video stream and the inter-frame hop count relationship in each GOP group are extracted. The fluctuation value of the sound source node range and the coding structure similarity of the current segmented video stream are calculated. If the fluctuation value of the sound source node range and the coding structure similarity both meet the first preset condition, the current segmented video stream is determined as the first video stream. Furthermore, by clustering each first video stream, the first video stream set is obtained.
[0027] For each segmented video stream in the segmented video stream set, the audio quality relationship between each sound source node with the same layout in the current segmented video stream is extracted, and if the audio quality relationship meets the second preset condition, the current segmented video stream is determined as the second video stream; further, by clustering each second video stream, the second video stream set is obtained.
[0028] The first and second preset conditions can be set according to the actual situation.
[0029] Optionally, the fluctuation value of the sound source node range in the current segmented video stream can be calculated in the following way: Using a target detection algorithm, the structured observation regions of all evaluators in each frame of the current segmented video stream are extracted, namely the "sound source node range" (usually a detection box). For each sound source node, the area change of the current sound source node in consecutive frames is tracked, and the area change rate of the sound source node is calculated. The area change rates of all sound source nodes are aggregated (e.g., the maximum value, average value, or variance is taken) to obtain the sound source node range fluctuation value of the current segmented video stream.
[0030] Optionally, the coding structure similarity of the current segmented video stream can be calculated in the following ways: Extract each GOP group from the current segmented video stream. For each GOP group, obtain the inter-frame hop count relationship of each B-frame in the current GOP group, represented by the total hop count TotalJump(b) (i.e., the sum of forward hops and backward hops). Based on the TotalJump(b) values of all B-frames in the current GOP group, construct the coding structure features of the current GOP group (such as the mean, variance, or distribution histogram of TotalJump(b)). By calculating the similarity of the coding structure features between each GOP group in the current segmented video stream (such as cosine similarity or correlation coefficient), the coding structure similarity of the current segmented video stream can be obtained.
[0031] The inter-frame hop count relationship (represented by the total hop count) can be determined in the following way: Let the starting I-frame (keyframe) of a GOP be located at sequence position i, and its first directly following P-frame be located at position p. Define the interval between the I-frame and the P-frame as D, calculated as: D = pi – 1. Therefore, the position of the P-frame can be represented as: p = i + D + 1. In this structure, the reference frame index of the P-frame is fixed to the starting I-frame of its GOP, i.e.: Ref(p) = pD - 1 = i.
[0032] For a B-frame at position b: its forward reference frame is the nearest I-frame or P-frame that precedes it, and its position is i; its backward reference frame is the nearest P-frame that follows it, and its position is p.
[0033] Based on the above reference frame, this B-frame is defined as follows: Forward jump count: bi-1; Back jump count: pb-1; Total Jump: This is the sum of forward jumps and backward jumps. The simplified core calculation formula is as follows: TotalJump(b)=(bi-1)+(pb-1)=pi-2; Encoding structure similarity is used to help determine the integrity and coherence of video data. The principle is that the inter-frame hop count reflects the regularity of video encoding. When the video stream encoding parameters are consistent and there is no abnormal processing, the inter-frame hop count of each GOP remains stable. If the video is corrupted due to post-production editing or abnormal transcoding, the hop count at the GOP boundaries will change abnormally. By calculating the consistency of the inter-frame hop count among the GOPs within the same segment of the video stream, encoding structure similarity can be obtained, thus allowing the selection of segmented video streams with coherent content and no abnormal editing.
[0034] S130. Determine the target video stream that belongs to both the first video stream set and the second video stream set.
[0035] Specifically, by using the intersection of the first and second video stream sets, video segments that meet the requirements for coding structure stability and may have abnormal audio activity are selected and used as target video streams for subsequent joint analysis.
[0036] S140. For any GOP group in the target video stream as the current GOP group, generate the action granularity of the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group.
[0037] In this embodiment, the evaluation content can be understood as the business data synchronized with the I-frames in the current GOP group in the evaluation system. This business data includes the document interface presented during the evaluation process, including but not limited to the bid opening summary table, technical clarification letter, business qualification certificate, price scoring table, and comprehensive conclusion report. The criticality of the evaluation content is used to characterize the importance level of the business data, and its function is to dynamically adjust the granularity of actions; the higher the criticality, the finer the action granularity, and the lower the criticality, the coarser the action granularity. The criticality is determined based on the screen proportion and the number of direct features of the business data; the larger the screen proportion and the more direct features, the higher the criticality.
[0038] Among them, action granularity refers to the detection accuracy of the behavioral characteristics of the evaluators during the video anomaly recognition process. The finer the action granularity, the higher the detection accuracy, the smaller the scope of the recognition object, and the higher the sampling frequency; the coarser the action granularity, the lower the detection accuracy, the larger the scope of the recognition object, and the lower the sampling frequency. For example, the more critical the evaluation content (such as the comprehensive scoring stage), the finer the detection, which can identify subtle actions such as flipping a phone or passing a note; the more routine the evaluation content (such as the document review stage), the coarser the detection, only identifying obvious anomalies such as leaving the post or obstructing the view, avoiding indiscriminate waste of computing power.
[0039] Specifically, for any GOP group in the target video stream as the current GOP group, first determine the evaluation content in the evaluation system corresponding to the I-frame (keyframe) in the current GOP group. Then, calculate the criticality of the evaluation content corresponding to the current GOP group. Based on the criticality, generate the action granularity adapted to the current GOP group to provide an analysis basis for subsequent anomaly detection based on the action granularity.
[0040] S150. Based on the sound source node layout of the current GOP group, the action granularity is adapted to obtain the target action granularity, and audio anomaly analysis is performed in combination with the spatial audio characteristics of the current GOP group.
[0041] In this embodiment, the sound source node layout refers to the spatial distribution and relative positional relationship of each sound source node in the video frame, which is a core visual feature reflecting the arrangement of people on site. For example, in the I-frame of the current GOP group, the sound source nodes are spatially distributed in a double-track pattern (two rows side by side), a single-sided pattern (a single row), or a closed-loop pattern (circular arrangement). The target action granularity is the detection accuracy suitable for the current scene after adapting the action granularity based on the sound source node layout of the current GOP group. The spatial audio features are the position coordinates and audio feature parameters of each sound source node in the three-dimensional sound field, used to locate abnormal sound sources and assist in the determination of audio anomalies.
[0042] Specifically, the sound source node layout of the current GOP group is first determined. Then, based on the sound source node layout, the action granularity is adaptively adjusted to obtain the target action granularity. Finally, combined with the spatial audio characteristics of the current GOP group, audio anomaly analysis is performed.
[0043] S160. Based on the results of the audio anomaly analysis and the target action granularity, perform anomaly identification on the current GOP group.
[0044] Specifically, based on the results of the audio anomaly analysis and the target action granularity, anomaly identification is performed on the current GOP group: if an audio anomaly is determined to exist, video anomaly identification is performed according to the target action granularity; if no audio anomaly is determined to exist, video anomaly identification is performed by combining the inter-frame explicit features and the target action granularity.
[0045] The present invention provides a technical solution that, according to the evaluation stage to which the video belongs, divides multiple evaluation video data into a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream; based on the fluctuation value of the sound source node range and the similarity of the coding structure of each segmented video stream, a first set of video streams is determined by clustering; based on the audio quality relationship of each sound source node in each segmented video stream, a second set of video streams is determined by clustering; a target video stream that simultaneously belongs to both the first and second set of video streams is determined; for any Group of Parts (GOPs) in the target video stream as the current GOP, an action granularity of the current GOP is generated based on the criticality of the evaluation content corresponding to the current GOP; based on the sound source node layout of the current GOP, the action granularity is adapted to obtain a target action granularity, and combined with the spatial audio features of the current GOP, audio anomaly analysis is performed; based on the results of the audio anomaly analysis and the target action granularity, anomaly identification is performed on the current GOP. First, based on the segmented multi-channel evaluation video data from the evaluation stage, a first video stream set is determined by combining the fluctuation values of the sound source node range and the similarity of the coding structure. Simultaneously, a second video stream set is determined by utilizing the audio quality relationships between sound source nodes. The target video stream is located through intersection filtering, and static segments are eliminated to reduce false alarms. Second, for the GOP group of the target video stream, action granularity is generated based on the criticality of the evaluation content, and adaptive adjustments are made according to the stability of the sound source node layout. Audio anomaly analysis is performed in conjunction with spatial audio features. Finally, anomaly identification is performed based on the analysis results and the target action granularity. This achieves synergy between fine analysis of key links and efficient filtering of routine links, thereby improving accuracy and reducing the false alarm rate.
[0046] In some embodiments, generating the action granularity of the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group includes: extracting features from the current GOP group using a preset feature extraction model to obtain initial action features; determining the evaluation content corresponding to the current GOP group using a preset mapping relationship between keyframes and evaluation content in the evaluation system; and scaling and adjusting the initial action features according to the criticality of the evaluation content to obtain the action granularity.
[0047] In this embodiment, the preset feature extraction model is a pre-trained analytical model used to extract behavioral features of evaluators from the current GOP group. For example, the preset feature extraction model can be a Two-Stream CNN.
[0048] The pre-defined mapping relationship between keyframes and evaluation content in the evaluation system refers to the pre-established temporal association rules between keyframes and evaluation content in the evaluation system. The process of establishing this mapping relationship is as follows: Find the corresponding I-frame in the set FI of all I-frames in the video stream. The timestamp of this frame satisfy Furthermore, it is the I-frame with the largest timestamp among all I-frames that meet this condition, meaning that this frame is the closest one to the I-frame that appeared before or at the same time as the evaluation content was generated. That keyframe: ; Through the above steps, at the time point (I-frame timing) and evaluation content A mapping relationship M is established between them: ; In subsequent use, for a given time point t, first find the evaluation content that is generated first after t. The time is Then find the time period corresponding to this evaluation content. Associated I-frame The evaluation content corresponding to this I-frame is For any point on the video timeline T, the evaluation content in the evaluation system at the corresponding time moment is found through the nearby I-frames.
[0049] Specifically, the initial motion features corresponding to the I-frame of the current GOP group are extracted using a preset feature extraction model; the mapping relationship between preset keyframes and evaluation content in the evaluation system is used to determine the evaluation content corresponding to the I-frame of the current GOP group in the evaluation system; then, the criticality of this evaluation content is calculated, and the calculation process is as follows: First, calculate the weighting of the length of the evaluation content. : ; Among them, SP is the initial image format value, which is 0.6 for half-frame and 1 for full-frame; LB is the number of gray value blocks in half-frame; LG is the number of non-gray value blocks. It's a comprehensive reflection of the image format and the proportion of key content. A shot with a full-frame sensor and a full view of the document will receive a much higher score than a shot with a half-frame sensor and only a small amount of document. value.
[0050] Secondly, according to the weight of length GP assesses the criticality of the evaluation content: ; Where S represents the number of initial motion features. GP is a comprehensive critical indicator; the GP value is higher when the image contains both important evaluation materials and complex motions.
[0051] Finally, based on the criticality of the evaluation content, the initial action features are scaled and adjusted to obtain the action granularity.
[0052] Optionally, the initial motion feature PAG can be scaled and adjusted in the following way: ; Among them, GP represents the criticality of the current GOP group's evaluation content; These represent the maximum and minimum values of the criticality of the evaluation content between the current GOP group and the adjacent GOP groups before and after it; when The higher the criticality level (GP), the greater its contribution to the granularity of actions, and the more closely the granularity of actions aligns with the importance of the evaluation content. Used to limit the range of granularity fluctuations, when The granularity of the action is smoothed by the criticality of the GOP groups before and after, avoiding drastic fluctuations in the granularity of the analysis caused by brief action noise.
[0053] By dynamically adjusting the granularity of actions, it achieves refined detection of key evaluation links and coarse-grained screening of non-critical links, effectively balancing detection accuracy and computational resource consumption, and improving the efficiency and accuracy of video anomaly recognition.
[0054] In some embodiments, the step of adapting the action granularity based on the sound source node layout of the current GOP group to obtain a target action granularity, and performing audio anomaly analysis in conjunction with the spatial audio features of the current GOP group, includes: calculating the stability of the sound source node layout based on the sound source node layout of the current GOP group; calculating the action spike granularity of the current GOP group in response to the stability being lower than a preset threshold, and determining the target action granularity based on the action spike granularity and the action granularity; determining the action granularity as the target action granularity in response to the stability being greater than or equal to the preset threshold; and performing audio anomaly analysis on the current GOP group based on the spatial audio features of the current GOP group and the target action granularity.
[0055] In this embodiment, the stability of the sound source node layout is a numerical indicator that quantifies the degree of change in the spatial distribution pattern of all sound source nodes within the current GOP group over time. Higher stability indicates a more fixed arrangement of evaluators and a more stable image; lower stability indicates more drastic changes in the movement or aggregation of evaluators. A preset threshold is a pre-set value used to determine whether the sound source node layout is stable. If the stability is lower than the preset threshold, the layout is considered unstable, and motion granularity needs to be calculated; if the stability is greater than or equal to the preset threshold, the layout is considered stable, and motion granularity is directly used. Motion granularity is a quantitative indicator of the precision of detecting sudden actions (such as suddenly taking out an item, quickly getting up, etc.) in the image, specifically for scenarios with unstable sound source node layouts. It can supplement the detection blind spots of motion granularity and enhance the ability to capture abnormal actions.
[0056] Specifically, extract the sound source node layout of the current GOP group and calculate the stability of the sound source node layout: Calculate the normalized coordinates of each sound source node relative to the central sound source node to form a feature vector. : ; in, The coordinates of the central sound source node are: Here are the coordinates of the first sound source node. The coordinates of the remaining nodes are represented in the same way. W and H are the width and height of the video frame, respectively.
[0057] Calculate the layout difference between consecutive frames To quantify the degree of drastic changes in layout: ; in, The L2 norm of a vector. The larger the value, the more drastic the change in the layout of the sound source nodes between two adjacent frames.
[0058] Within the current GOP group A convergent evaluation is performed to obtain the stability score of the current GOP group. : ; in, For the current GOP group The mean of the value is used to measure the average level of change; the smaller the value, the more stable the system. For the current GOP group Standard deviation is used to measure the degree of fluctuation; the smaller the value, the more stable the system. α and β are preset adjustment coefficients.
[0059] By combining the action granularity PAG, the stability Be of the sound source node layout is calculated: ; Wherein, γ is the preset adjustment coefficient.
[0060] In response to stability, Be is greater than or equal to a preset threshold. At this point, the sound source node layout is stable, and there is no need to calculate the granularity of the action thrust. The action granularity PAG can be directly determined as the target action granularity.
[0061] In response to stability, Be is less than a preset threshold. If the sound source node layout is unstable at this point, then the action spike granularity SPIKE is calculated based on the stability Be: ; Where M is the total number of spike events identified in the current GOP group. A spike event is an abnormal feature that is higher than the median in the initial action features. Consecutive abnormal feature frames can be merged into a spike event. Let be the number of frames lasting the i-th thrust event; The time weighting factor is the inverse of the time span of the start or end phase of an action, and N is the normalization factor, such as the total number of frames T or the average feature change of the entire sequence, to make SPIKE comparable.
[0062] The action granularity PAG is updated based on the action thrust granularity SPIKE to obtain the target action granularity: PAG = SPIKE + PAG.
[0063] Extract the spatial audio stream Rd that is temporally aligned with the current GOP group. Calculate the comprehensive stability index Be of this spatial audio stream Rd, which integrates characteristics such as short-term burstiness, spectral stability, and source localization consistency of the signal.
[0064] If the overall stability index Be is less than or equal to the preset threshold T_stable1, then it is determined that there is a risk of audio anomalies within the current GOP group time period: Source separation is performed on the spatial audio stream Rd to obtain independent audio channels corresponding to each source node; short-time framing is performed on each independent audio channel, the segment-level stability of each audio segment is calculated, suspicious audio segments with stability lower than the sub-threshold E_stable are filtered out, and their spatial location coordinates Pdi(x,y,z) are determined; all suspicious segments and their coordinates constitute the suspicious event set {Pdi} in the current GOP group.
[0065] For each suspicious audio segment, its acoustic features (such as spectrum, Mel frequency cepstral coefficients, etc.) are extracted and fused with the corresponding spatial location coordinates Pdi to form a spatial audio feature vector.
[0066] Based on the safety level and target action granularity (PAG) of the current GOP group's evaluation stage, the parameters of the audio anomaly classification model (such as support vector machine (SVM)) are dynamically configured; the spatial audio feature vectors are input into the configured classification model for judgment, and the analysis results of audio anomaly analysis are output.
[0067] By dynamically adjusting the granularity of target actions through the stability of sound source node layout, it achieves adaptive switching between routine detection in stable scenes and refined spike detection in unstable scenes. Combined with spatial audio features, it performs cross-modal anomaly analysis, effectively improving the accuracy and efficiency of identifying abnormal behavior in complex bidding environments.
[0068] In some embodiments, the step of identifying anomalies in the current GOP group based on the results of the audio anomaly analysis and the target action granularity includes: if the results of the audio anomaly analysis indicate that there is an audio anomaly in the current GOP group, then based on the target action granularity, a preset video understanding large model is used to detect the current GOP group to obtain an abnormal behavior identification result; if the results of the audio anomaly analysis indicate that there is no audio anomaly in the current GOP group, then based on the video features of the current GOP group and the target action granularity, a preset base video analysis model is used to determine abnormal frames to obtain the location of the abnormal scene and the cause of the anomaly, wherein the video features include color block fluctuation curvature and inter-frame brightness gradient values.
[0069] In this embodiment, the preset video understanding big model can be understood as a pre-trained deep learning model, possessing the ability to understand video image semantics, extract behavioral features, and reason about complex behaviors. It can perform full-scale abnormal behavior identification and structured judgment on video segments based on the granularity of target actions; the video understanding big model used in this embodiment is specifically the video-SALMONN-o1 model. The preset base video analysis model can be understood as a pre-trained basic video analysis model, possessing the core capabilities of video frame feature extraction, accurate abnormal frame localization, and abnormal cause determination. It focuses on fine-grained frame-level analysis of video images and can combine color block fluctuation curvature, inter-frame brightness gradient values, and target action granularity to complete frame-level anomaly detection; the base video analysis model used in this embodiment is specifically the Gen MAC model. Color block undulation curvature is a feature index that quantifies abnormal changes in color block artifacts in video images. It is used to capture color block artifacts caused by improper data compression or data transmission errors due to abnormal changes in the image during video encoding. In low bitrate transmission or abnormal transcoding scenarios, unnatural blocky effects appear in color areas that should have smooth transitions in the image. By calculating the size, distribution, and edge undulation curvature of color blocks in video frames, abnormal areas in the video caused by personnel leaving their posts, video interruptions, carrying prohibited items, illegal gatherings, changing seats, image occlusion, and frequent hand gesture interactions can be detected, providing spatial feature basis for frame-level anomaly judgment. Inter-frame brightness gradient value is a feature index that quantifies the temporal changes in the brightness relationship between adjacent video frames. By calculating the difference in brightness gradient between corresponding pixels in adjacent frames, it captures the abrupt changes in brightness and abnormal temporal changes in light and shadow caused by actions in the image. This feature, combined with color block undulation curvature, achieves complementarity between the spatial features and temporal change features of the video image, improving the accuracy of the base video analysis model in identifying abnormal frames.
[0070] Specifically, if the audio anomaly analysis indicates that there is an audio anomaly in the current GOP group, then based on the system's built-in video understanding model (such as the video-SALMONN-o1 model), according to the target action granularity, the video understanding model is guided to detect and infer the full range of behaviors in the current GOP group, such as personnel leaving their posts, video interruption, carrying prohibited items, illegal gathering, exchanging seats, screen occlusion, and frequent gesture interactions, to obtain the abnormal behavior identification results; if the audio anomaly analysis indicates that there is no audio anomaly in the current GOP group, the color block fluctuation curvature and inter-frame brightness gradient values of each frame in the current GOP group are extracted. Then, the color block fluctuation curvature, inter-frame brightness gradient values, the current GOP group, and the target action granularity are input into the preset base video analysis model (such as the Gen MAC model) to obtain the location of the abnormal screen and the cause of the anomaly.
[0071] By leveraging audio anomaly analysis results, intelligent separation is achieved between the large-scale video understanding model and the base video analysis model. This allows the large-scale model to focus on in-depth reasoning in complex scenarios with audio anomalies, while the base model handles efficient frame-level detection in scenarios without audio anomalies, thus achieving synergistic improvement in both computing resource optimization and detection accuracy.
[0072] In some embodiments, the step of clustering to determine the first video stream set based on the source node range fluctuation value and coding structure similarity of each segmented video stream includes: determining the segmented video streams that meet preset conditions as the first video stream; wherein, the preset conditions include the source node range fluctuation value being greater than the source node range fluctuation threshold and the coding structure similarity being greater than the jump relationship similarity threshold; and determining the first video stream set based on each first video stream.
[0073] In this embodiment, the preset conditions can be understood as pre-set criteria for selecting video segments with substantial action content that have not been abnormally edited. The sound source node range fluctuation threshold is a quantitative threshold used to determine whether the sound source node range has undergone substantial changes. When the sound source node range fluctuation value is greater than this threshold, it indicates that there are substantial activities such as personnel position changes, gatherings, or dispersals in the current segmented video stream, rather than a static or fixed camera position. The jump relationship similarity threshold is a quantitative threshold used to determine the consistency of the coding structure between each GOP group within the video stream. When the coding structure similarity is greater than this threshold, it indicates that the video coding structure is stable and coherent; when the similarity is lower than this threshold, it indicates that there are coding parameter jumps or structural breaks at the GOP boundaries, suggesting possible post-editing, abnormal transcoding, or malicious interruption.
[0074] Specifically, for each segmented video stream, the segmented video streams whose sound source node range fluctuation value is greater than the sound source node range fluctuation threshold and whose coding structure similarity is greater than the jump relationship similarity threshold are identified as the first video stream; then, each first video stream is clustered to determine the first video stream set.
[0075] By using a dual screening method based on the fluctuation value of the sound source node range and the similarity of the coding structure, video clips that are static, lack substantial behavior, or have undergone post-editing or abnormal transcoding are effectively excluded. High-quality video data with coherent content and substantial behavior are retained, providing a reliable data foundation for subsequent granular generation of actions and anomaly recognition.
[0076] In some embodiments, the audio quality relationships are represented by a quality matrix, which is determined based on loudness, brightness, specific frequencies, and anomalous audio tags. The anomalous audio tags are used to identify audio anomalies in the segmented video streams. The step of clustering to determine a second video stream set based on the audio quality relationships of each sound source node in each segmented video stream includes: determining the segmented video streams with a number of sound source nodes whose anomalous audio tags are not empty (greater than a preset node number threshold) as the second video streams; and determining the second video stream set based on each second video stream.
[0077] Specifically, the audio quality relationships of each sound source node in the current segmented video stream can be calculated in the following way: The audio data of each sound source node in the current segmented video stream is extracted and processed frame by frame. For each sound source node, the audio feature parameters of the current sound source node are extracted frame by frame, including loudness (L), luminance (C), and intensity (THD) of a specific frequency component. Then, the average loudness, average luminance, and average intensity of the specific frequency component of the current sound source node are calculated across all frames in the current segmented video stream, thereby forming an audio feature vector representing the current sound source node.
[0078] Based on the audio feature vectors of all sound source nodes, the global median (i.e., the median of the average feature values for each sound source node) is calculated for loudness, brightness, and intensity of specific frequency components. For each sound source node, the average value of each feature in the current sound source node's audio feature vector is compared with the corresponding global median, and the number of features for the current sound source node that exceed the corresponding global median is counted. It is worth noting that if the number of features is less than a preset threshold for the number of abnormal features, the abnormal audio label SFER corresponding to the current sound source node is set to empty; otherwise, the abnormal audio label SFER is set to the number of features, i.e., not empty.
[0079] Therefore, a quality matrix F=[L,C,THD,SFER] is constructed, where L is the average loudness, C is the average brightness, THD is the average intensity of a specific frequency component, and SFER is the abnormal audio tag.
[0080] If SFER is not empty, it is initially determined that the current sound source node has abnormal audio behavior. Then, the number of sound source nodes with non-empty SFER is counted. If the total number exceeds a preset node number threshold, the current segmented video stream is determined as the second video stream that needs to be evaluated for audiovisual integration.
[0081] By using multidimensional representation of the quality matrix and quantification of abnormal audio tags, the system can accurately locate audio anomalies and perform statistical screening of groups, effectively improving the detection efficiency of illegal audio risks in bidding scenarios.
[0082] Example 2 Figure 3This is a schematic diagram of an anomaly detection device provided in Embodiment 2 of the present invention. Figure 3 As shown, the device includes: The video segmentation module 21 is used to segment multiple bidding video data according to the bidding stage to which the video belongs, to obtain a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream. The set determination module 22 is used to cluster and determine the first video stream set based on the fluctuation value of the sound source node range and the similarity of the coding structure of each segment video stream, and to cluster and determine the second video stream set based on the audio quality relationship of each sound source node in each segment video stream. The target video stream determination module 23 is used to determine a target video stream that simultaneously belongs to both the first video stream set and the second video stream set; The granularity determination module 24 is used to take any GOP group in the target video stream as the current GOP group and generate the action granularity of the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group. The audio analysis module 25 is used to adapt the action granularity based on the sound source node layout of the current GOP group to obtain the target action granularity, and to perform audio anomaly analysis in combination with the spatial audio features of the current GOP group. The anomaly identification module 26 is used to identify anomalies in the current GOP group based on the results of the audio anomaly analysis and the target action granularity.
[0083] The technical solution provided in Embodiment 2 of this invention firstly involves segmenting multiple video data streams during the evaluation phase, determining a first video stream set by combining the fluctuation values of the sound source node range and the similarity of the coding structure, and simultaneously determining a second video stream set by utilizing the audio quality relationship between sound source nodes. The target video stream is located through intersection filtering, and static segments are eliminated to reduce false alarms. Secondly, for the GOP group of the target video stream, action granularity is generated based on the criticality of the evaluation content, and adaptation adjustments are made according to the stability of the sound source node layout. Audio anomaly analysis is performed by combining spatial audio features, and finally, anomaly identification is performed based on the analysis results and the target action granularity. This achieves synergy between refined analysis of key links and efficient filtering of routine links, thereby improving accuracy and reducing the false alarm rate.
[0084] Optionally, the granularity determination module 24 includes: The feature extraction unit is used to extract features from the current GOP group using a preset feature extraction model to obtain initial action features; The content determination unit is used to determine the evaluation content corresponding to the current GOP group by using the mapping relationship between preset keyframes and evaluation content in the evaluation system; The granularity determination unit is used to scale and adjust the initial action features according to the criticality of the evaluation content to obtain the action granularity.
[0085] Optional, the audio analysis module 25 includes: The stability determination unit is used to calculate the stability of the sound source node layout based on the current GOP group's sound source node layout. The first target determination unit is used to calculate the action thrust granularity of the current GOP group in response to the stability being lower than a preset threshold, and determine the target action granularity based on the action thrust granularity and the action granularity. The second target determination unit determines the action granularity as the target action granularity in response to the stability being greater than or equal to a preset threshold. The audio analysis unit is used to perform audio anomaly analysis on the current GOP group based on the spatial audio characteristics of the current GOP group and the target action granularity.
[0086] Optionally, the anomaly detection module 26 includes: The first result determination unit is used to detect the current GOP group based on the target action granularity and using a preset video understanding large model if the result of audio anomaly analysis indicates that there is an audio anomaly in the current GOP group, thereby obtaining the abnormal behavior identification result. The second result determination unit is used to determine abnormal frames based on the video features of the current GOP group and the target action granularity if the result of the audio anomaly analysis is that there is no audio anomaly in the current GOP group. The abnormal frame location and the cause of the anomaly are then determined by using a preset base video analysis model. The video features include color block fluctuation curvature and inter-frame brightness gradient values.
[0087] Optionally, the set determination module 22 includes: The first video stream determination unit is used to determine the segmented video stream that meets the preset conditions from each segmented video stream as the first video stream; wherein, the preset conditions include the sound source node range fluctuation value being greater than the sound source node range fluctuation threshold, and the coding structure similarity being greater than the jump relationship similarity threshold. The first set determination unit is used to determine the first video stream set based on each first video stream.
[0088] Optionally, the audio quality relationship is represented by a quality matrix, which is determined based on loudness, brightness, specific frequencies and anomalous audio tags, and the anomalous audio tags are used to identify audio anomalous features in the segmented video stream; Optionally, the set determination module 22 includes: The second video stream determination unit is used to determine the segmented video streams in which the number of sound source nodes with non-empty abnormal audio tags is greater than a preset node number threshold as the second video stream. The second set determination unit is used to determine the second video stream set based on each second video stream.
[0089] Optionally, the video may be part of at least one of the following bidding stages: capital verification, bid splitting, bid reading, review, negotiation, and comprehensive conclusion.
[0090] Optionally, the sound source node corresponds to the bid evaluation personnel at the bid evaluation site; the node layout is one of the following: double-track, single-sided, and closed-loop. The double-track layout includes a layout in which bid evaluation personnel are arranged in two parallel rows, the single-sided layout includes a layout in which bid evaluation personnel are arranged in a single row, and the closed-loop layout includes a layout in which bid evaluation personnel are arranged in a ring.
[0091] The anomaly identification device provided in the embodiments of the present invention can execute the anomaly identification method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0092] Example 3 Figure 4 This is a schematic diagram of an electronic device provided in Embodiment 3 of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0093] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0094] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0095] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as anomaly detection methods.
[0096] In some embodiments, the anomaly detection method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or mounted on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the anomaly detection method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the anomaly detection method by any other suitable means (e.g., by means of firmware).
[0097] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0098] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0099] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0100] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0101] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0102] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0103] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0104] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
[0105] This invention also provides a computer program product, including a computer program and / or instructions, which, when executed by a processor, implements the anomaly identification method provided in any embodiment of this application.
[0106] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0107] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. An anomaly identification method, characterized in that, include: Based on the bidding stage to which the video belongs, the bidding video data from multiple channels is divided to obtain a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream; Based on the fluctuation values of the sound source node range and the similarity of the coding structure of each segment of the video stream, the first video stream set is determined by clustering. Based on the audio quality relationship of each sound source node in each segment of the video stream, the second video stream set is determined by clustering. Determine the target video stream that simultaneously belongs to both the first video stream set and the second video stream set; For any GOP group in the target video stream as the current GOP group, the action granularity of the current GOP group is generated based on the criticality of the evaluation content corresponding to the current GOP group. Based on the sound source node layout of the current GOP group, the action granularity is adaptively adjusted to obtain the target action granularity, and audio anomaly analysis is performed in combination with the spatial audio characteristics of the current GOP group. Based on the results of the audio anomaly analysis and the target action granularity, anomaly identification is performed on the current GOP group.
2. The method according to claim 1, characterized in that, The generation of action granularity for the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group includes: Using a preset feature extraction model, features are extracted from the current GOP group to obtain initial action features; By utilizing the mapping relationship between preset keyframes and the evaluation content in the evaluation system, the evaluation content corresponding to the current GOP group is determined; Based on the criticality of the evaluation content, the initial action features are scaled and adjusted to obtain the action granularity.
3. The method according to claim 1, characterized in that, The process involves adapting the action granularity based on the sound source node layout of the current GOP group to obtain the target action granularity, and then performing audio anomaly analysis in conjunction with the spatial audio characteristics of the current GOP group, including: Calculate the stability of the sound source node layout based on the current GOP group's sound source node layout; In response to the stability being lower than a preset threshold, the action thrust granularity of the current GOP group is calculated, and the target action granularity is determined based on the action thrust granularity and the action granularity. In response to the stability being greater than or equal to a preset threshold, the action granularity is determined as the target action granularity; Based on the spatial audio features of the current GOP group and the target action granularity, audio anomaly analysis is performed on the current GOP group.
4. The method according to claim 1, characterized in that, The anomaly identification of the current GOP group based on the results of the audio anomaly analysis and the target action granularity includes: If the audio anomaly analysis results indicate that there is an audio anomaly in the current GOP group, then based on the target action granularity, the current GOP group is detected using a preset video understanding large model to obtain the abnormal behavior identification result. If the audio anomaly analysis results indicate that there are no audio anomalies in the current GOP group, then based on the video features of the current GOP group and the target action granularity, anomaly frame determination is performed using a preset base video analysis model to obtain the location and cause of the anomaly. The video features include color block fluctuation curvature and inter-frame brightness gradient values.
5. The method according to claim 1, characterized in that, The first video stream set is determined by clustering based on the fluctuation values of the sound source nodes and the similarity of the coding structure of each segmented video stream, including: The video stream segment that meets the preset conditions in each video stream segment is determined as the first video stream; wherein, the preset conditions include the sound source node range fluctuation value being greater than the sound source node range fluctuation threshold, and the coding structure similarity being greater than the jump relationship similarity threshold. Based on each first video stream, determine the first video stream set.
6. The method according to claim 1, characterized in that, The audio quality relationships are represented by a quality matrix, which is determined based on loudness, brightness, specific frequencies, and anomalous audio tags. The anomalous audio tags are used to identify audio anomalous features in the segmented video stream. The method of clustering to determine the second video stream set based on the audio quality relationships of each sound source node in each segmented video stream includes: The segmented video stream in which the number of sound source nodes with non-empty abnormal audio tags is greater than a preset node number threshold is identified as the second video stream. Based on each second video stream, determine the set of second video streams.
7. The method according to claim 1, characterized in that, The video belongs to at least one of the following bidding stages: capital verification, bid splitting, bid reading, review, negotiation, and comprehensive conclusion.
8. The method according to claim 1, characterized in that, The sound source node corresponds to the bid evaluation personnel at the bid evaluation site; the node layout is one of the following: double-track, single-sided, and closed-loop. The double-track layout includes a layout in which bid evaluation personnel are arranged in two parallel rows; the single-sided layout includes a layout in which bid evaluation personnel are arranged in a single row; and the closed-loop layout includes a layout in which bid evaluation personnel are arranged in a ring.
9. An anomaly detection device, characterized in that, include: The video segmentation module is used to segment multiple bidding video data according to the bidding stage to which the video belongs, to obtain a set of segmented video streams, wherein the set of segmented video streams includes at least one segmented video stream; The set determination module is used to cluster and determine the first video stream set based on the fluctuation value of the sound source node range and the similarity of the coding structure of each segment video stream, and to cluster and determine the second video stream set based on the audio quality relationship of each sound source node in each segment video stream. The target video stream determination module is used to determine a target video stream that simultaneously belongs to both the first video stream set and the second video stream set. The granularity determination module is used to take any GOP group in the target video stream as the current GOP group and generate the action granularity of the current GOP group based on the criticality of the evaluation content corresponding to the current GOP group. The audio analysis module is used to adapt the action granularity based on the sound source node layout of the current GOP group to obtain the target action granularity, and to perform audio anomaly analysis in combination with the spatial audio characteristics of the current GOP group. An anomaly identification module is used to identify anomalies in the current GOP group based on the results of the audio anomaly analysis and the target action granularity.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the anomaly identification method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the anomaly identification method according to any one of claims 1-8.
12. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the anomaly identification method according to any one of claims 1-8.