Method for extracting key audio-video

CN120730094BActive Publication Date: 2026-08-21XIAOSHAN DISTRICT BRANCH OF HANGZHOU PUBLIC SECURITY BUREAU +2
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510951458.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2026-08-21
Estimated Expiration
2045-07-10

AI Technical Summary

Technical Problem

这一过程需要执法人员投入大量的时间和精力,逐一查看、筛选和整理数据,严重影响了工作效率,增加了执法人员的工作负担

Benefits of technology

1、实现一个报案号对应多个执法人员,实现小组模式下的多人执法的自动绑定;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120730094B_ABST
    Figure CN120730094B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of audio and video intelligent processing, in particular to a key audio and video extraction method, which comprises the following steps: based on case number information, one or more alarm number information and audio and video corresponding to the case number information are acquired and added to a processing list; taking the case receiving time in the case number information as a node, the audio and video in the processing list are segmented to obtain corresponding video streams; all the video streams are decoded and pretreated to generate continuous frame images and audio tracks; the case number information bound with the video streams is used for estimating a police alarm starting node; through two ways of explicit keyword detection and implicit signal analysis, an event trigger node and / or an event termination node are generated; based on the event trigger node and the event termination node, corresponding video streams are processed through timestamps, and complete key audio and video segments are extracted; the application has the advantages of fast and accurate realization of key audio and video extraction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of intelligent audio and video processing, and in particular to a method for extracting key audio and video data. Background Technology

[0002] In the ongoing process of building a society governed by the rule of law, the standardization, transparency, and fairness of law enforcement work are receiving increasing attention. As a crucial tool for recording law enforcement processes, body cameras play a key role in ensuring fairness and protecting the legitimate rights and interests of law enforcement personnel and the public. When performing various tasks, law enforcement officers use body cameras to record audio and video, accurately recreating the enforcement scene and providing solid evidence for subsequent dispute resolution and case review.

[0003] However, existing law enforcement recorders have significant limitations in practical use. Currently, most law enforcement recorders rely on law enforcement officers manually turning on the camera to record after arriving at the designated location. This method has many drawbacks. On the one hand, law enforcement officers often face complex, changeable, and urgent situations when carrying out tasks, and their attention is highly focused on the law enforcement matter itself. They are very likely to forget to turn on the camera due to negligence, resulting in the loss of crucial records of the law enforcement process. For example, at the scene of an emergency, law enforcement officers need to respond quickly to various situations and may not have time to turn on the recorder, making that part of the law enforcement process unrecorded and causing many problems for subsequent work. On the other hand, even if law enforcement officers remember to turn on the recorder, they may not be familiar with its operation or have time constraints, failing to turn on the device in a timely and accurate manner, missing the recording of key law enforcement steps, and affecting the integrity of the law enforcement record.

[0004] To address the aforementioned issues, our company has innovatively developed a solution combining a law enforcement recorder and a smart cabinet. When the recorder is placed in the smart cabinet, the two interact via specific connection methods, such as physical interface connection and wireless communication connection. In this way, the system can automatically trigger the camera to turn on the moment the recorder is retrieved from the smart cabinet, automatically and continuously recording the entire process from retrieval to return to the cabinet. This effectively avoids omissions that may occur due to manual operation by law enforcement personnel, ensuring that the entire process of major and sensitive incidents is effectively recorded from beginning to end, providing strong support for law enforcement work and reliable evidence for supervision and inspection.

[0005] Meanwhile, the solution also features convenient data processing capabilities. After law enforcement officers return the body camera to the smart cabinet upon completion of their task, the cabinet connects to the camera via USB cable, charging it while automatically acquiring the audio and video data stored within. This design greatly facilitates data collection and organization, improving work efficiency.

[0006] However, this innovative solution has also revealed new problems in practical applications. Because the body camera records the entire process from retrieval to return, the resulting audio and video content is longer or more abundant. A significant proportion of this massive amount of audio and video data is invalid, such as irrelevant footage of officers traveling from the smart cabinet to the enforcement scene, or non-critical scenes during breaks in enforcement. After completing their daily enforcement work, officers need to accurately extract the relevant portions of the data from this vast amount of audio and video data to meet the requirements of case organization, archiving, and subsequent review, and then accurately bind them to their badge numbers and corresponding case numbers. This process requires officers to invest a significant amount of time and energy, reviewing, filtering, and organizing the data one by one, severely impacting work efficiency and increasing their workload. Summary of the Invention

[0007] In order to quickly and accurately extract key audio and video data for binding with law enforcement personnel and case report numbers, this application provides a method for extracting key audio and video data.

[0008] This application provides a method for extracting key audio and video data, employing the following technical solution: A method for extracting key audio and video data, comprising the following steps: Based on the case report number, obtain the corresponding multiple police siren information, and add the audio and video of the law enforcement recorder bound to the police siren information to the processing list; and take the case acceptance time in the case report number information as the node, remove the police siren information corresponding to the audio and video that is empty under the corresponding time node in the processing list, and remove the police siren information corresponding to the audio and video that has been bound in the processing list. Using the case acceptance time in the case report number information as a node, the audio and video in the processing list are segmented to obtain the corresponding video stream; Decode and preprocess all video streams to generate consecutive frame images and separate the audio track from the video stream; An enforcement path is generated based on the reporting location and the address of the law enforcement recorder at the time of receiving the case, which are bound to the case number information. The travel time is calculated based on the enforcement path, and the start node of the police response is estimated by the case acceptance time and the travel time. Using the alarm start node as the starting time, all audio is subjected to explicit keyword detection. When an explicit keyword is detected, an event trigger node and / or an event termination node are generated. When no event trigger node and / or event termination node are generated, the latent signal is obtained based on the frame image and audio track, and the event trigger node and / or event termination node are generated based on the multimodal combination of the latent signal. Retrieve the bound audio and video segments of the removed siren information. If neither the event trigger node nor the event termination node falls into the bound audio and video segments, then add the audio and video of the law enforcement recorder bound to the siren information back to the processing list. Based on the corresponding event trigger node and event termination node generated for each audio and video in the processing list, the corresponding video stream is processed by timestamp, so that each audio and video generates a complete key audio and video segment.

[0009] In one embodiment: the nodes that generate event triggering nodes and / or event termination nodes based on the multimodal union of latent signals include: Based on the preset weight of each latent signal, the time period in which the total weight is greater than the threshold within the preset duration is obtained. Based on the time order, the start time of the first time period is selected as the start time node, and the acquisition time of the last latent signal in the last time period is selected as the end time node. Based on the start time node, and combined with the backtracking duration of the implicit signal within the first time period, the event trigger node is generated after correction. Based on the end time node, and combined with the delay duration of the implicit signal in the last time period, the event termination node is generated after correction.

[0010] In one embodiment: when only the event triggering node is not generated, the acquisition of implicit signals is stopped after a time period is obtained; when only the event termination node is not generated, only the end time node is obtained.

[0011] In one embodiment: the step of obtaining time periods with a total weight greater than a threshold within a preset duration, and selecting the start time of the first time period as the start time node based on the time sequence, specifically includes: After acquiring the latent signal, based on the time node where the latent signal is located as the starting node, acquire all latent signals within a preset time period; The total weight is obtained by summing the weights of all implicit signals within a preset time period. When the total weight is greater than the threshold, the starting node is taken as the start time node; When the total weight is less than the threshold, the next adjacent latent signal is obtained based on the time sequence, and the total weight is recalculated based on the time node where the latent signal is located.

[0012] In one embodiment: the step of obtaining time periods within a preset duration where the total weight is greater than a threshold, and selecting the acquisition time of the last latent signal in the last time period as the end time node based on the time sequence, specifically includes: After acquiring all latent signals, the time node of the last latent signal is used as the termination node to acquire all latent signals within a preset time period. The total weight is obtained by summing the weights of all implicit signals within a preset time period. When the total weight is greater than the threshold, the termination node will be used as the end time node. When the total weight is less than the threshold, the previous adjacent latent signal is obtained based on the time sequence, and the time node where the latent signal is located is used as the termination node to recalculate the total weight.

[0013] In one embodiment: the step of generating the event trigger node after correction based on the start time node and the backtracking duration of the implicit signal within the first time period specifically includes: Based on the implicit signals included in the first time period, obtain the preset backtracking duration corresponding to the implicit signals; The backtracking time is calculated based on the weights corresponding to the latent signals and the preset backtracking time. Using the start time node as the anchor point, the event trigger node is generated by extending backwards by the backtracking time.

[0014] In one embodiment: the step of generating an event termination node based on the end time node and the delay duration of the implicit signal in the last time period specifically includes: Based on the implicit signals included in the last time period, obtain the preset delay duration corresponding to the implicit signals; The delay time is calculated based on the weights corresponding to the latent signals and the preset delay time. Using the end time node as the anchor point, the event termination node is generated by extending it backwards by the subsequent delay.

[0015] In one embodiment: the implicit signals include scene switching signals, semantic signals, and environmental change signals; Specifically, the step of obtaining the latent signal based on the frame image and audio track includes: When an enforcement certificate or information matching the complainant information in the case number is first detected in a frame image, a scene switching signal is generated to indicate the start of enforcement; when no information matching the complainant information in the case number appears in a frame image for N seconds, a scene switching signal is generated to indicate the end of enforcement. Detect hidden keywords in audio and output semantic signals indicating the start or end of law enforcement based on the type of hidden keywords. When an audio change from being dominated by other sounds to being dominated by human voices is detected, an environmental change signal is generated to indicate the start of law enforcement; when an audio change from being dominated by human voices to being dominated by other sounds is detected, an environmental change signal is generated to indicate the end of law enforcement.

[0016] In one embodiment: the cause of the report in the report number information is input into a preset model to generate an estimated processing time, and then the end point of the emergency response is calculated based on the emergency response start point.

[0017] In one embodiment: the video stream contains GPS data, and the latent signal further includes a location mutation signal; the method for obtaining the location mutation signal specifically includes the following steps: Extract GPS data from the video stream; Based on the time sequence analysis of GPS data latitude and longitude changes, when the speed is detected to become zero and remain so for a first preset time, a location change signal is generated to indicate the start of law enforcement; when the speed is detected to suddenly increase from zero and not return to zero within a second preset time, a location change signal is generated to indicate the end of law enforcement.

[0018] In summary, this application has the following beneficial effects: 1. Enable one case number to correspond to multiple law enforcement officers, and achieve automatic binding of multiple law enforcement personnel in team mode; 2. By segmenting lengthy audio and video recordings, audio and video recordings prior to the case acceptance time can be directly removed, directly shortening the time required to extract audio and video recordings corresponding to each case number and effectively reducing the processing cycle. 3. By estimating the start point of the emergency response, the length of audio and video data to be processed can be further reduced. During processing, explicit keyword analysis of audio is prioritized. By extracting explicit keywords spoken by law enforcement officers at the start and end of enforcement, event trigger and termination nodes are directly generated. Then, the corresponding audio and video data can be extracted based on these two nodes. Unlike video processing, audio can quickly eliminate a large amount of invalid data. Therefore, this method is the fastest, most accurate, and simplest, and can greatly reduce the computational requirements. 4. When law enforcement officers forget to say explicit keywords or they are covered by noise and fail to be identified, the implicit signal multimodal joint judgment method is used to generate event trigger nodes and event termination nodes. The multimodal joint judgment method ensures the accuracy of key audio and video extraction. 5. It can perform joint processing of audio and video based on multiple alarm siren information, effectively improving the accuracy of key audio and video extraction. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the key audio and video extraction method in this embodiment; Figure 2 This is a flowchart illustrating the multimodal joint judgment method in this embodiment. Detailed Implementation

[0020] The present application will be further described in detail below with reference to the accompanying drawings.

[0021] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.

[0022] It should be noted that the descriptions of these embodiments are for the purpose of aiding understanding the present invention, but do not constitute a limitation thereof. Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0023] In the description of this application, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0024] In the description of this application, the terms "one embodiment," "some embodiments," "illustrative embodiment," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples.

[0025] A method for extracting key audio and video, such as Figure 1 As shown, it includes the following steps S100. Based on the case report number information, obtain the corresponding multiple siren information, add the audio and video of the law enforcement recorder bound to the siren information to the processing list; and take the case acceptance time in the case report number information as the node, remove the siren information corresponding to the audio and video that is empty under the corresponding time node in the processing list, and remove the siren information corresponding to the audio and video that is already bound in the processing list.

[0026] The report number information includes the report number, reporting location, reporter information, time of acceptance, address of the body camera at the time of acceptance, and reason for the report. The reporter information includes the reporter's name, gender, address, house number, photo, and license plate number. The police identification information mainly includes the law enforcement officer's information, such as the officer's name and identification number. In this step, the address of the body camera at the time of acceptance is generated by the body camera's built-in location function and is simultaneously linked to the report number information.

[0027] When a case report number corresponds to multiple police numbers, it may indicate that multiple law enforcement officers are involved. In this case, it's necessary to bind the audio and video recordings of all officers to the case report number. Furthermore, in cases involving multiple officers, some officers may be unavailable due to other commitments. Therefore, it's necessary to remove the audio and video recordings of these officers to reduce subsequent workload.

[0028] The removal process primarily addresses two scenarios: First, there are no audio / video records after the case acceptance time in the case report information, meaning the content at that time point is empty. Second, law enforcement officers are not conducting enforcement actions due to other ongoing activities. In this case, removal is performed using pre-linked audio / video recordings, meaning the case acceptance time is included in the pre-linked audio / video recordings. However, since the case acceptance time is not the enforcement time in this situation, this removal is only temporary and requires further verification.

[0029] Furthermore, to avoid conflicts, case report numbers must be bound sequentially based on time order. However, to improve binding speed and enable simultaneous binding of multiple case report numbers, when multiple case report numbers are identified under a single alarm number, these alarm number information can be bound sequentially based on time order.

[0030] S200. Using the case acceptance time in the case report number information as a node, segment the audio and video in the processing list to obtain the corresponding video stream.

[0031] Depending on the recording mode of the law enforcement recorder, more and more such devices now automatically limit the length of each video during recording. For example, without stopping recording, they create a separate video segment every 30 minutes and store it. Therefore, the video stream mentioned above can be a video segment formed by dividing a complete long video, or it can be a video set formed by dividing several videos.

[0032] For video sets, if the case acceptance time is within a certain video segment, the video segment can be divided based on the case acceptance time as a node, or the entire video segment can be directly included in the video stream.

[0033] In this step, after audio and video segmentation is completed, the video stream is bound to the case report number information, enabling the system to divide a long audio and video into multiple independent video streams. This allows the system to process multiple video streams simultaneously, shortening the processing cycle. Furthermore, if law enforcement personnel need to adjust key extracted audio and video segments later, they can directly retrieve the corresponding video stream for processing, reducing the workload of manual processing.

[0034] S300 performs decoding preprocessing on all video streams to generate consecutive frame images and separates the audio track from the video stream.

[0035] In this step, the generated frame images must be aligned with the timestamps of the audio track. Algorithms such as Weiner filtering and spectral subtraction are used to remove background noise from the audio, including sounds like the friction of law enforcement recorders, wind noise, vehicle noise, and crowd noise, thus improving speech clarity. Simultaneously, to reduce unnecessary computation, silent segments in the audio can be removed through silence detection. For example, based on an energy threshold (short-term energy + zero-crossing rate), audio segments with an energy below -30dB and lasting for more than 200ms are considered silent segments.

[0036] S400 generates an enforcement path based on the reporting location and the address of the enforcement recorder at the time of receiving the case, which are bound to the case number information. It also calculates the travel time based on the enforcement path and estimates the start node of the police response based on the case acceptance time and travel time.

[0037] The start time of the police response is calculated as: the time of receiving the case plus the travel time. For example, if the time of receiving the case is "20XX-04-24 10:00:00" and the travel time is 30 minutes, then the start time of the police response is "20XX-04-24 10:30:00".

[0038] This step can be achieved by directly calling the map routing API. In addition, to improve the accuracy of the calculation, the mode of transportation can be bound to the siren information to generate more accurate travel time based on different modes of transportation.

[0039] Additionally, it should be noted that since the calculation of the police response start point is performed after the fact, the enforcement route and travel time do not need to be generated based on real-time traffic data. However, to further improve accuracy, the road conditions corresponding to the enforcement time can be simulated using algorithms to generate the data.

[0040] S500: Detect explicit keywords in the audio using the alarm start node as the starting time. When an explicit keyword is detected, generate an event trigger node and / or an event termination node.

[0041] In this embodiment, explicit keywords refer to a word or phrase agreed upon in advance that law enforcement officers need to say before and after enforcement, such as "This enforcement begins," "Enforcement begins," "Enforcement begins for case number XX," "This enforcement ends," "Enforcement ends," "Enforcement ends for case number XX," etc. To ensure accuracy and reduce human error, multiple explicit keywords can be set to form a thesaurus.

[0042] When S600, event trigger node and / or event termination node are not generated, latent signals are obtained based on frame images and audio tracks, and event trigger nodes and / or event termination nodes are generated based on the multimodal combination of latent signals.

[0043] In addition, in this embodiment, the implicit signals include scene switching signals, semantic signals, and environmental change signals. Each of these signals includes signals representing the start and end of law enforcement, respectively. Signals related to the start of law enforcement are used to generate event trigger nodes, while signals related to the end of law enforcement are used to generate event termination nodes. Different scene switching signals, different semantic signals, and different environmental change signals are all preset with different weights, and all weights are required to be less than 1.

[0044] Specifically, the steps for obtaining latent signals based on frame images and audio tracks include: Using object detection algorithms such as YOLO, a scene transition signal is generated to indicate the start of law enforcement when an enforcement certificate or information matching the complainant's information in the case report number is detected for the first time in a frame image. Conversely, a scene transition signal is generated to indicate the end of law enforcement when no information matching the complainant's information appears in the frame image for N seconds. The information matching the complainant's information in the case report number mainly refers to facial features, license plates, house numbers, and the complainant's identification documents.

[0045] The system detects hidden keywords in the audio and outputs semantic signals indicating the start or end of law enforcement based on the type of hidden keywords. The hidden keywords in this step mainly refer to words and phrases that law enforcement officers are likely to use during the start and end of law enforcement phases, such as "checking ID," "reporter information," and "you can leave." Furthermore, the hidden keyword lexicon can be customized and expanded.

[0046] When an audio change from being dominated by other sounds to being dominated by human voice is detected, an environmental change signal is generated to indicate the start of law enforcement; when an audio change from being dominated by human voice to being dominated by other sounds is detected, an environmental change signal is generated to indicate the end of law enforcement. In this step, the dominance type of the sound can be determined by setting a sliding window and then statistically analyzing the ratio of human voice frames to ambient sound frames within the window. Simultaneously, to eliminate the influence of transient fluctuations, a hysteresis threshold can be used for determination. That is, after a change in dominance, such as from being dominated by other sounds to being dominated by human voice, it is only considered a change from being dominated by other sounds to being dominated by human voice if subsequent sliding windows consistently detect that human voice remains dominant.

[0047] In another embodiment, the cause of the report in the report number information is input into a preset model to generate an estimated processing time, and then the end time of the response is calculated based on the start time node. The start time node is generated in step S400. The start time node and the end time node also represent the commencement and termination of law enforcement, respectively.

[0048] In another embodiment, the video stream contains GPS data, and the implicit signal also includes a location mutation signal, wherein the method for obtaining the location mutation signal includes the following steps: Extract GPS data from the video stream; Based on the time sequence analysis of GPS data latitude and longitude changes, when the speed is detected to become zero and remain so for a first preset time, a location change signal is generated to indicate the start of law enforcement; when the speed is detected to suddenly increase from zero and not return to zero within a second preset time, a location change signal is generated to indicate the end of law enforcement.

[0049] Specifically, the straight-line distance between two adjacent GPS points is calculated using the Haversine formula. Instantaneous velocity can be calculated using the straight-line distance and the time interval between the two adjacent GPS points. The corresponding positioning change signal can be generated by judging the instantaneous velocity.

[0050] The Haversine formula is as follows:

[0051] In the formula, d is the straight-line distance between two adjacent GPS points (in meters), R is the Earth's radius (approximately 6371 kilometers), ∅1 and ∅2 are the latitudes of the two points, and λ1 and λ2 are the longitudes of the two points, in radians.

[0052] S700: Obtain the bound audio and video segments of the removed siren information. If none of the generated event trigger nodes and event termination nodes fall into the bound audio and video segments, then re-add the audio and video of the law enforcement recorder bound to the siren information to the processing list.

[0053] Based on the different audio and video processing in the processing list, each audio and video segment generates its corresponding event trigger node and event termination node. This is mainly because different law enforcement officers may arrive at the scene at different times. Therefore, in the above identification process, the earliest and latest time nodes should be used as the standard. Thus, it is sufficient to ensure that all event trigger nodes and event termination nodes do not fall into the bound audio and video segments.

[0054] S800: Based on the corresponding event trigger node and event termination node generated for each audio and video in the processing list, the corresponding video stream is processed by timestamp, so that each audio and video generates a complete key audio and video segment.

[0055] In this step, the corresponding event triggering node and event termination node generated for each audio and video in the list are processed to generate a complete key audio and video segment for each audio and video.

[0056] In another embodiment, such as Figure 2 As shown, the multimodal joint judgment method specifically includes: S610. Based on the preset weight of each latent signal, obtain the time period within the preset duration where the total weight is greater than the threshold. Based on the time order, select the start time of the first time period as the start time node and select the acquisition time of the last latent signal in the last time period as the end time node.

[0057] In this step, to reduce computational load, when only the event trigger node has not been generated, the acquisition of implicit signals is stopped after a certain time period. Conversely, when only the event termination node has not been generated, only the end time node needs to be acquired.

[0058] The specific steps of obtaining the time periods with a total weight greater than a threshold within a preset duration, and selecting the start time of the first time period as the starting time node based on the time sequence, include: After obtaining the latent signal, based on the time node where the latent signal is located as the starting node, all latent signals within a preset time period are obtained.

[0059] The total weight is obtained by summing the weights of all implicit signals within a preset time period.

[0060] When the total weight is greater than the threshold, the starting node is used as the start time node.

[0061] When the total weight is less than the threshold, the next adjacent latent signal is obtained based on the time sequence, and the time node where the latent signal is located is used as the starting node to recalculate the total weight.

[0062] Correspondingly, the steps of obtaining the time period within the preset duration where the total weight is greater than the threshold, and selecting the acquisition time of the last latent signal in the last time period as the end time node based on the time sequence, specifically include: After acquiring all latent signals, the time node of the last latent signal is used as the termination node to acquire all latent signals within a preset time period.

[0063] The total weight is obtained by summing the weights of all implicit signals within a preset time period.

[0064] When the total weight is greater than the threshold, the termination node will be used as the end time node.

[0065] When the total weight is less than the threshold, the previous adjacent latent signal is obtained based on the time sequence, and the time node where the latent signal is located is used as the termination node to recalculate the total weight.

[0066] In the above steps, the set threshold is greater than the weight of any latent signal to avoid making a judgment directly based on a single latent signal.

[0067] S620. Based on the start time node and combined with the backtracking duration of the implicit signal within the first time period, the event trigger node is generated after correction.

[0068] This step specifically includes: Based on the implicit signals included in the first time period, obtain the preset backtracking duration corresponding to the implicit signals.

[0069] The backtracking time is calculated based on the weights corresponding to the latent signals and the preset backtracking time.

[0070] Using the start time node as the anchor point, the event trigger node is generated by extending backwards by the backtracking time.

[0071] S630. Based on the end time node and combined with the delay duration of the implicit signal in the last time period, the event termination node is generated after correction.

[0072] This step specifically includes: Based on the implicit signal included in the last time period, obtain the preset delay time corresponding to the implicit signal.

[0073] The delay time is calculated based on the weights corresponding to the latent signals and the preset delay time.

[0074] Using the end time node as the anchor point, the event termination node is generated by extending it backwards by the subsequent delay.

[0075] The implicit signals include scene switching signals, semantic signals, environmental change signals, process node signals, and location change signals. The scene switching signals, semantic signals, environmental change signals, location change signals, and process node signals all contain two types: one indicates the start of law enforcement, and the other indicates the end of law enforcement.

[0076] In steps S620 and S630, depending on the type and nature of the implicit signal, the start of enforcement corresponds to a preset backtracking time, and the end of enforcement corresponds to a preset delay time. The preset backtracking time and preset delay time are different for each implicit signal.

[0077] For example, if the detected latent signals are scene switching signals and semantic signals, the preset backtracking time for a scene switching signal indicating the start of law enforcement is 10 minutes with a weight of A1, and the preset delay time for a scene switching signal indicating the end of law enforcement is 5 minutes with a weight of A2; the preset backtracking time for a semantic signal indicating the start of law enforcement is 15 minutes with a weight of B1, and the preset delay time for a semantic signal indicating the end of law enforcement is 5 minutes with a weight of B2. Then, the calculated backtracking time = A1 * 10 minutes + B1 * 15 minutes, and the calculated delay time = A2 * 5 minutes + B2 * 5 minutes.

[0078] The event trigger node = start time node - calculation backtracking duration, and the event termination node = end time node + calculation delay duration.

[0079] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A method for extracting key audio and video, characterized in that, Includes the following steps: Based on the case report number, obtain the corresponding multiple police siren information, add the audio and video of the law enforcement recorder bound to the police siren information to the processing list; and use the case acceptance time in the case report number information as a node to remove the police siren information corresponding to audio and video that is empty under the corresponding time node in the processing list, as well as remove the police siren information corresponding to audio and video that is already bound in the processing list. Using the case acceptance time in the case report number information as a node, the audio and video in the processing list are segmented to obtain the corresponding video stream; Decode and preprocess all video streams to generate consecutive frame images and separate the audio track from the video stream; An enforcement path is generated based on the reporting location and the address of the law enforcement recorder at the time of receiving the case, which are bound to the case number information. The travel time is calculated based on the enforcement path, and the start node of the police response is estimated by the case acceptance time and the travel time. The audio is subjected to explicit keyword detection with the alarm start node as the starting time. When an explicit keyword is detected, an event trigger node and / or event termination node are generated. When no event trigger node and / or event termination node are generated, the latent signal is obtained based on the frame image and audio track, and the event trigger node and / or event termination node are generated based on the multimodal combination of the latent signal. Get the bound audio and video of the removed siren information. If none of the generated event trigger nodes and event termination nodes fall into the bound audio and video segments, then add the audio and video of the law enforcement recorder bound to the siren information back to the processing list. Based on the corresponding event trigger node and event termination node generated for each audio and video in the processing list, the corresponding video stream is processed by timestamp, so that each audio and video generates a complete key audio and video segment.

2. The method for extracting key audio and video according to claim 1, characterized in that, The nodes that generate event trigger nodes and / or event termination nodes based on the multimodal joint of latent signals include: Based on the preset weight of each latent signal, the time period in which the total weight is greater than the threshold within the preset duration is obtained. Based on the time order, the start time of the first time period is selected as the start time node, and the acquisition time of the last latent signal in the last time period is selected as the end time node. Based on the start time node, and combined with the backtracking duration of the implicit signal within the first time period, the event trigger node is generated after correction. Based on the end time node, and combined with the delay duration of the implicit signal in the last time period, the event termination node is generated after correction.

3. The method for extracting key audio and video according to claim 2, characterized in that: When only the event trigger node has not been generated, the acquisition of implicit signals is stopped after a time period is obtained; when only the event termination node has not been generated, only the end time node is obtained.

4. The method for extracting key audio and video according to claim 2, characterized in that, The steps for obtaining the time periods with a total weight greater than a threshold within a preset duration, and selecting the start time of the first time period as the starting time node based on chronological order, specifically include: After acquiring the latent signal, based on the time node where the latent signal is located as the starting node, acquire all latent signals within a preset time period; The total weight is obtained by summing the weights of all implicit signals within a preset time period. When the total weight is greater than the threshold, the starting node is taken as the start time node; When the total weight is less than the threshold, the next adjacent latent signal is obtained based on the time sequence, and the total weight is recalculated based on the time node where the latent signal is located.

5. The method for extracting key audio and video according to claim 2, characterized in that, The steps of obtaining the time periods within a preset duration where the total weight is greater than a threshold, and selecting the acquisition time of the last latent signal in the last time period as the end time node based on the time sequence, specifically include: After acquiring all latent signals, the time node of the last latent signal is used as the termination node to acquire all latent signals within a preset time period. The total weight is obtained by summing the weights of all implicit signals within a preset time period. When the total weight is greater than the threshold, the termination node will be used as the end time node. When the total weight is less than the threshold, the previous adjacent latent signal is obtained based on the time sequence, and the time node where the latent signal is located is used as the termination node to recalculate the total weight.

6. The method for extracting key audio and video according to claim 2, characterized in that, The step of generating an event trigger node based on the start time node and the backtracking duration of the implicit signal within the first time period specifically includes: Based on the implicit signals included in the first time period, obtain the preset backtracking duration corresponding to the implicit signals; The backtracking time is calculated based on the weights corresponding to the latent signals and the preset backtracking time. Using the start time node as the anchor point, the event trigger node is generated by extending backwards by the backtracking time.

7. The method for extracting key audio and video according to claim 2, characterized in that, The step of generating an event termination node based on the end time node and the delay duration of the implicit signal in the last time period specifically includes: Based on the implicit signals included in the last time period, obtain the preset delay duration corresponding to the implicit signals; The delay time is calculated based on the weights corresponding to the latent signals and the preset delay time. Using the end time node as the anchor point, the event termination node is generated by extending it backwards by the subsequent delay.

8. The method for extracting key audio and video according to any one of claims 2-7, characterized in that: The implicit signals include scene switching signals, semantic signals, and environmental change signals; The steps for obtaining latent signals based on frame images and audio tracks specifically include: When an enforcement certificate or information matching the complainant information in the case number is first detected in a frame image, a scene switching signal is generated to indicate the start of enforcement; when no information matching the complainant information in the case number appears in a frame image for N seconds, a scene switching signal is generated to indicate the end of enforcement. Detect hidden keywords in audio and output semantic signals indicating the start or end of law enforcement based on the type of hidden keywords. When an audio change from being dominated by other sounds to being dominated by human voices is detected, an environmental change signal is generated to indicate the start of law enforcement; when an audio change from being dominated by human voices to being dominated by other sounds is detected, an environmental change signal is generated to indicate the end of law enforcement.

9. The method for extracting key audio and video according to claim 8, characterized in that: After inputting the cause of the report from the report number into the preset model, the estimated processing time is generated, and then the end time of the emergency response is calculated based on the start node of the emergency response.

10. The method for extracting key audio and video according to claim 8, characterized in that: The video stream contains GPS data, and the implicit signal also includes a location change signal. The method for obtaining the location change signal specifically includes the following steps: Extract GPS data from the video stream; Based on the time sequence analysis of GPS data latitude and longitude changes, when the speed is detected to become zero and remain so for a first preset time, a location change signal is generated to indicate the start of law enforcement; when the speed is detected to suddenly increase from zero and not return to zero within a second preset time, a location change signal is generated to indicate the end of law enforcement.

Citation Information

Patent Citations

  • Information automatic correlation method and system

    CN103530359A

  • Video processing method, video processing equipment and storage medium

    CN113259761A

  • Video processing method and device, equipment and storage medium

    CN118200464A

  • Live video key point marking method and device, equipment and storage medium

    CN119071520A

  • Method and system for segmenting and transmitting on-demand live-action video in real-time

    US20120219271A1