Video data processing method and device, electronic equipment and medium
By extracting key event segments from video data from multiple cameras in industrial scenarios and merging them, a large-scale risk identification model is used to identify safety risks. This solves the problem of the difficulty in efficiently identifying safety risks in industrial scenarios in existing technologies, and achieves more efficient risk monitoring and identification.
Patent Information
- Application Number
- CN202511261517.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-02
AI Technical Summary
Existing technologies struggle to efficiently identify and promptly alert to security risks when processing surveillance video data collected from multiple cameras in industrial settings.
By extracting key event segments related to security risks from the video data of each camera, merging and processing them, and inputting them into a large risk identification model, the model identifies the risks associated with the acquisition location of each target video data and comprehensively judges the security risks of the target scene.
It enables efficient risk monitoring and identification in complex industrial production scenarios, improving the accuracy and efficiency of risk identification.
Smart Images

Figure CN121053584A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more particularly to the fields of large-scale modeling, data processing, and safety production management, specifically to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing video data. Background Technology
[0002] In industrial safety management, the storage and processing of 24-hour uninterrupted monitoring video data streams collected by a large number of cameras are usually involved. Therefore, the processing system has high requirements for safety risk identification in such scenarios.
[0003] Currently, large-scale visual models can be used to process multiple surveillance video data collected by multiple cameras set up in industrial scenarios to identify whether any safety risks have occurred during the production process in the industrial scenario and to provide timely alerts.
[0004] The methods described in this section are not necessarily methods that had been previously conceived or adopted. Unless otherwise specified, no method described in this section should be assumed to be prior art simply because it is included in this section. Similarly, unless otherwise specified, the issues mentioned in this section should not be considered to be accepted in any prior art. Summary of the Invention
[0005] This disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for processing video data.
[0006] According to one aspect of this disclosure, a method for processing video data is provided, comprising: acquiring multiple surveillance video data for a target scene; extracting video segments associated with key events from the multiple surveillance video data to obtain multiple video segments, wherein the key events indicate that the acquisition location of the surveillance video data may have security risks; merging the multiple video segments to obtain at least one target video data; processing the at least one target video data using a risk identification big data model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether there is a security risk and the type of security risk present at the acquisition location associated with the corresponding target video data; and determining a security risk monitoring result for the target scene based on the at least one initial identification result.
[0007] According to another aspect of this disclosure, a model training method is provided, comprising: acquiring a training dataset, wherein the training dataset includes multiple training video data and annotation results for the multiple training video data, the annotation results indicating the security risk type of the acquisition location associated with the multiple training video data; and training an initial large model using the training dataset to obtain a risk identification large model, wherein the risk identification large model is used to identify whether there is a security risk associated with the acquisition location associated with the target video data and the type of security risk present.
[0008] According to another aspect of this disclosure, a video data processing apparatus is provided, comprising: a first module configured to acquire multiple surveillance video data corresponding to multiple acquisition locations of a target scene; a second module configured to extract video segments associated with key events from the multiple surveillance video data to obtain multiple video segments, wherein the key events indicate that the acquisition locations of the surveillance video data may have security risks; a third module configured to merge the multiple video segments to obtain at least one target video data; a fourth module configured to process the at least one target video data using a risk identification big data model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether there is a security risk at the acquisition location associated with the corresponding target video data and the type of security risk present; and a fifth module configured to determine a security risk monitoring result for the target scene based on the at least one initial identification result.
[0009] According to another aspect of this disclosure, a model training apparatus is provided, comprising: a sixth module configured to acquire a training dataset, wherein the training dataset includes multiple training video data and annotation results for the multiple training video data, the annotation results indicating the security risk type of the acquisition location associated with the multiple training video data; and a seventh module configured to train an initial large model using the training dataset to obtain a risk identification large model, wherein the risk identification large model is used to identify whether there is a security risk associated with the acquisition location of the target video data and the type of security risk present.
[0010] According to another aspect of this disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the methods described above.
[0011] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform the above-described method.
[0012] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described method.
[0013] According to one or more embodiments of this disclosure, a video data processing method is provided. For safety production management scenarios involving multiple cameras, video clips of key events potentially associated with safety risks are first extracted from the corresponding monitoring video data based on the acquisition location of each camera. Then, the extracted video clips are merged and processed to obtain at least one target video data point, which is input into a risk identification model. The risk identification model identifies the risk identification result at the acquisition location associated with each target video data point, and the safety risk monitoring result of the target scenario is comprehensively judged based on the risk identification results from multiple acquisition locations. Therefore, risk monitoring and identification can be better performed for complex industrial production scenarios.
[0014] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0015] The accompanying drawings exemplify embodiments and form part of the specification, serving together with the textual description to explain exemplary implementations of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals refer to similar but not necessarily identical elements.
[0016] Figure 1 This is a schematic diagram illustrating an example system in which various methods described herein may be implemented according to exemplary embodiments; Figure 2 A flowchart illustrating a method for processing video data according to an embodiment of the present disclosure is shown; Figure 3 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 4 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 5 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 6 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 7 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 8 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown; Figure 9 A flowchart of a model training method according to an embodiment of the present disclosure is shown; Figure 10 A structural block diagram of a video data processing apparatus according to an embodiment of the present disclosure is shown; Figure 11 A structural block diagram of a model training apparatus according to an embodiment of the present disclosure is shown; and Figure 12 A structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present disclosure is shown. Detailed Implementation
[0017] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0018] In this disclosure, unless otherwise stated, the use of terms such as "first," "second," etc., to describe various elements is not intended to limit the positional, temporal, or importance relationships of these elements; such terms are merely used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of that element, while in other cases, based on the context, they may refer to different instances.
[0019] The terminology used in the description of the various examples described in this disclosure is for the purpose of describing particular examples only and is not intended to be limiting. Unless the context explicitly indicates otherwise, an element may be one or more unless the number of elements is specifically limited. Furthermore, the term "and / or" as used in this disclosure covers any one of the listed items and all possible combinations thereof.
[0020] In related technologies, large-scale visual models can be used to process multiple monitoring video data collected by multiple cameras set up in industrial scenarios to identify whether any safety risks have occurred in the production process of industrial scenarios and to provide timely alerts. To address the aforementioned issues, this disclosure provides a video data processing method. For safety production management scenarios involving multiple cameras, it first extracts video clips of key events potentially associated with safety risks from the corresponding monitoring video data based on the acquisition location of each camera. Then, it merges and processes these extracted video clips to obtain at least one target video data point, which is input into a large-scale risk identification model. This model identifies the risk at each acquisition location associated with the target video data, and the safety risk monitoring results for the target scenario are comprehensively judged based on the risk identification results from multiple acquisition locations. Therefore, it enables better risk monitoring and identification in complex industrial production scenarios.
[0021] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0022] Figure 1 A schematic diagram of an exemplary system 100 in which the various methods and apparatus described herein can be implemented according to embodiments of this disclosure is shown. Reference Figure 1 The system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to execute one or more applications.
[0023] In embodiments of this disclosure, server 120 may run one or more services or software applications that enable the execution of methods for processing video data.
[0024] In some embodiments, server 120 may also provide other services or software applications that may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, such as to users of client devices 101, 102, 103, 104, 105 and / or 106 under a Software as a Service (SaaS) model.
[0025] exist Figure 1 In the configuration shown, server 120 may include one or more components that implement the functions performed by server 120. These components may include software components, hardware components, or combinations thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 can sequentially interact with server 120 using one or more client applications to utilize the services provided by these components. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1This is an example of a system for implementing the video data processing methods described herein, and is not intended to be limiting.
[0026] Users can use client devices 101, 102, 103, 104, 105, and / or 106 to execute document generation methods. The client devices can provide an interface that allows users to interact with the client devices. The client devices can also output information to the user through this interface. Although... Figure 1 Only six client devices are described, but those skilled in the art will understand that this disclosure can support any number of client devices.
[0027] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, gaming systems, thin clients, various messaging devices, sensors, or other sensing devices. These computer devices can run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux or Linux-like operating systems (such as Google Chrome OS); or include various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablets, personal digital assistants (PDAs), etc. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices, internet-enabled gaming devices, etc. Client devices are capable of executing various applications, such as various internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0028] Network 110 can be any type of network well known to those skilled in the art, and can support data communication using any of a variety of available protocols (including but not limited to TCP / IP, SNA, IPX, etc.). By way of example only, one or more networks 110 can be a local area network (LAN), an Ethernet-based network, a token ring network, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0029] Server 120 may include one or more general-purpose computers, special-purpose server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.
[0030] The computing unit in server 120 can run one or more operating systems, including any of the aforementioned operating systems and any commercially available server operating system. Server 120 can also run any of a variety of additional server applications and / or middleware applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0031] In some implementations, server 120 may include one or more applications to analyze and merge data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and 106. Server 120 may also include one or more applications to display data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and 106.
[0032] In some implementations, server 120 can be a server for a distributed system or a server integrated with blockchain. Server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. A cloud server is a host product in the cloud computing service system, designed to address the shortcomings of traditional physical hosts and Virtual Private Server (VPS) services, such as high management difficulty and weak business scalability.
[0033] System 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as text files and video files. Databases 130 may reside in various locations. For example, a database used by server 120 may be local to server 120, or it may be located remotely to server 120 and may communicate with server 120 via a network-based or dedicated connection. Databases 130 may be of different types. In some embodiments, the database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from and from the databases in response to commands.
[0034] In some embodiments, one or more of the databases 130 may also be used by an application to store application data. The databases used by the application may be of different types, such as key-value stores, object stores, or regular stores supported by a file system.
[0035] Figure 1 The system 100 can be configured and operated in various ways to enable the application of the various methods and apparatus described in this disclosure.
[0036] Figure 2 A flowchart illustrating a method for processing video data according to an embodiment of the present disclosure is shown.
[0037] like Figure 2 As shown, the video data processing method 200 includes: Step 210: Obtain multiple surveillance video data for the target scene; Step 220: Extract video clips associated with key events from multiple surveillance video data to obtain multiple video clips, where the key events indicate that the location where the surveillance video data was collected may have security risks; Step 230: Merge multiple video segments to obtain at least one target video data; Step 240: Process at least one target video data using a large-scale risk identification model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether there is a security risk associated with the acquisition location of the corresponding target video data and the type of security risk present; and Step 250: Determine the security risk monitoring results for the target scenario based on at least one initial identification result.
[0038] Therefore, for safety production management scenarios involving multiple cameras, video clips of key events potentially associated with safety risks are first extracted from the corresponding monitoring video data based on the acquisition location of each camera. Then, the extracted video clips are merged and processed to obtain at least one target video data point, which is input into a large-scale risk identification model. This model then identifies the risk at the acquisition location associated with each target video data point, and comprehensively judges the safety risk monitoring results of the target scenario based on the risk identification results from multiple acquisition locations. This allows for better risk monitoring and identification in complex industrial production scenarios.
[0039] In step 210, the target scenario can be, for example, an industrial production scenario that requires safety risk monitoring and management, such as mining, power, oil and gas, water conservancy, chemical industry, and water affairs.
[0040] For example, since the target scenarios mentioned above usually involve a large area, it is necessary to set up multiple cameras in multiple locations to fully monitor the entire production process. Based on this, multiple monitoring video data can be collected by multiple cameras at multiple collection locations.
[0041] In step 210, according to some embodiments, the video surveillance data is video stream data. Since video stream data needs to be processed continuously, it places higher demands on the risk identification and processing system compared to ordinary video file data. The method described above, which first extracts video segments associated with key events, merges them, and then performs comprehensive identification, can achieve better processing results for video stream data.
[0042] In step 220, the operation of extracting video segments associated with key events may be performed separately for each of the multiple monitoring video data to obtain multiple video segments.
[0043] In step 220, the key events at different acquisition locations may be the same or different. For example, for a control room that must be staffed by on-duty personnel, a key event may be the detection of no one in the monitoring screen related to the control room; for some locations where personnel are prohibited from entering, a key event may be the detection of someone in the monitoring screen related to that location.
[0044] In addition, critical events can also include, for example, the detection of people smoking, gatherings of people, people not wearing specific work equipment, gas / liquid leaks, and equipment malfunctions at the corresponding data collection location.
[0045] Figure 3 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0046] According to some embodiments, such as Figure 3As shown, step 220, "extracting video clips associated with key events from multiple surveillance video data to obtain multiple video clips," includes: Step 310: Based on the corresponding acquisition locations of multiple surveillance video data, determine the event recognition model, wherein the event recognition model is used to identify key events; and Step 320: Use an event recognition model to process multiple surveillance video data to obtain multiple video clips.
[0047] Since different acquisition locations may involve different key events, a specific event recognition model can be used to extract video segments based on the specific acquisition location of the surveillance video data. This reduces the amount of data that the large risk recognition model needs to process later, and better balances the accuracy of the risk recognition results with the computing power requirements in the processing.
[0048] In step 310, the event recognition model can be a smaller visual model with lower recognition accuracy than the risk recognition model, thereby simplifying the model structure and processing. In one example, for each piece of surveillance video data from multiple surveillance video datasets, an event recognition model can be determined based on the corresponding acquisition location of that surveillance video data to identify key events at that acquisition location.
[0049] Figure 4 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0050] According to some embodiments, such as Figure 4 As shown, step 220 includes: Step 410: Determine multiple processing tasks corresponding to multiple monitoring video data for extracting video segments associated with the key events from the monitoring video data; Step 420: Determine the task sequence for the multiple monitoring video data in the thread pool based on the acquisition locations of the multiple monitoring video data corresponding to the multiple processing tasks; Step 430: Establish a task pool based on the task sequence to process multiple processing tasks, resulting in multiple video clips; and Step 440: Store the multiple video clips into their respective storage spaces.
[0051] In industrial scenarios with multiple cameras, it is often necessary to process surveillance video data captured by at least some of the cameras in parallel. Distributed storage can be used to store multiple surveillance video data sets, and a task pool can be introduced for resource scheduling and management. This task pool-based "pre-creation-reuse-dynamic management" model better enables complex resource scheduling and management of multiple surveillance video data sets on top of distributed storage, thereby improving resource scheduling efficiency during processing.
[0052] In step 410, each processing task corresponds to surveillance video data, and each processing task includes extracting video segments associated with key events from the surveillance video data corresponding to the processing task. Then, in step 420, a task sequence can be determined based on the respective acquisition locations of multiple surveillance video data.
[0053] In steps 410 and 420, the completed processing tasks are added to the thread pool's message queue to wait for scheduling of corresponding resources for processing. For example, the maximum number of threads in the thread pool can be limited to avoid creating too many threads and causing memory exhaustion in high-concurrency scenarios.
[0054] In step 430, for the currently pending task in the task sequence, a corresponding task pool can be created based on its specific task content to execute the task.
[0055] In step 440, by distributing and storing multiple video segments, dynamic expansion and contraction can be easily achieved, and the corresponding video segments can be quickly and concurrently invoked during subsequent processing, demonstrating good performance in high-concurrency processing scenarios. For example, the above-mentioned distributed storage can be implemented using a distributed storage system based on the minio (a distributed storage architecture system) architecture, thereby enabling parallel transmission across multiple nodes based on minio, effectively shortening data transmission time.
[0056] In one example, MySQL (a database type) can be used to store records, allowing tasks to be resent to the message queue after a task fails in the task pool. This enables features such as resume from interruption, error retries, and uploading in weak network conditions, improving fault tolerance during data transmission.
[0057] Figure 5 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0058] According to some embodiments, such as Figure 5 As shown, step 430 includes: Step 510: In response to determining that at least one main view task is included among multiple processing tasks, at least one corresponding task pool is established for at least one main view task, wherein the main view task indicates that the acquisition location of the corresponding monitoring video data is the main view location. Step 520: Based on at least one task pool, process at least one main-view task in parallel; and Step 530: In response to determining that a completed main-view task exists, obtain the auxiliary-view task associated with the completed main-view task from multiple processing tasks, wherein the auxiliary-view task indicates that the acquisition location of the corresponding monitoring video data is the auxiliary-view location; and Step 540: Reuse the task pool corresponding to the completed main-view task to process auxiliary-view tasks.
[0059] For each acquisition location, multiple cameras can be set up to capture surveillance video data from different angles to prevent blind spots. Specifically, one or more main-view cameras (e.g., shooting from a frontal view) can be set up for each acquisition location, along with one or more auxiliary-view cameras (e.g., shooting from a different angle than the main-view cameras) to provide richer visual information for subsequent risk identification.
[0060] Therefore, by establishing a task pool to prioritize and process the main-view tasks of the surveillance video data collected by the main-view camera in parallel, and then for each task pool, after processing the corresponding main-view tasks, the task pool is directly reused to process the corresponding auxiliary-view tasks of the surveillance video data collected by the auxiliary-view camera based on the same processing logic, which effectively reduces the overhead required to frequently create new task pools and improves resource utilization and system response speed.
[0061] In step 510, each task pool has specific task processing content and processing logic for extracting key events from the collection location of the corresponding main perspective task.
[0062] In step 520, by processing all main-view tasks in parallel, video clip extraction of key events at various acquisition locations of the target scene can be processed more completely with minimal computing resources, reducing the waiting time required before merging multiple video clips.
[0063] In steps 530 and 540, the auxiliary view task and the main view task target the same acquisition location but different acquisition angles. Therefore, the same task pool can be reused for processing to reduce the CPU (Central Processing Unit) and memory consumed by frequently creating and destroying thread resources.
[0064] In steps 530 and 540, for each task pool in at least one task pool, in response to determining that the main view task in the task pool is completed, the auxiliary view task associated with the completed main view task is obtained from multiple processing tasks, and then the task pool is reused to process the auxiliary view task.
[0065] In step 230, video segments can be filtered and merged based on the acquisition time and / or acquisition location.
[0066] Figure 6 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0067] According to some embodiments, such as Figure 6 As shown, step 230 includes: Step 610: Classify the multiple video segments based on their acquisition information to obtain at least one set of segments, wherein the acquisition information includes at least one of the acquisition location and acquisition time of the multiple video segments; and Step 620: Merge multiple video segments based on at least one segment set to obtain at least one target video data corresponding to at least one segment set.
[0068] Therefore, by selecting and merging video segments with high correlation based on the acquisition location and time, the merged target video data can include more complementary information, thus making the final recognition result more accurate.
[0069] In step 610, according to some embodiments, in response to the acquisition information including the acquisition locations of multiple video segments, the distance between the acquisition locations corresponding to any two video segments in the classified segment set is less than a distance threshold; and in response to the acquisition information including the acquisition time of multiple video segments, the time difference between the acquisition times corresponding to any two video segments in the classified segment set is less than a time threshold.
[0070] For example, surveillance video data collected by cameras that are close to each other have a high degree of correlation. Therefore, when the collected information includes the location of the collection, video clips in the same specific area can be merged so that the merged target video data can better characterize whether there is a security risk in that specific area.
[0071] For example, in a mine's transport corridor, video clips from all the cameras located in the corridor can be merged so that a large model can identify safety risks based on the complete monitoring content of the entire corridor.
[0072] For example, surveillance video data collected at close times have a high degree of correlation. Therefore, when the collected information includes the collection time, video clips collected at the same time can be merged to better characterize whether there are security risks within that time period.
[0073] For example, when devices at different acquisition locations that are far apart need to be shut down simultaneously, the video segments corresponding to the shutdown times can be merged so that the merged target video data can better represent whether there is a security risk that the devices have not been shut down simultaneously at the current moment.
[0074] In step 620, for each segment set in at least one segment set, the video segments in that segment set can be merged to obtain the target video data corresponding to that segment set, thereby obtaining at least one target video data corresponding to at least one segment set.
[0075] It is understood that the above merging and filtering process is for illustrative purposes only and is not intended to limit the scope. In this example, video clips can also be filtered and merged based on both capture time and capture location, which will not be elaborated upon further.
[0076] In step 240, the risk identification big model includes a visual big model, which is used to identify at least one target video data based on the input to obtain an initial identification result.
[0077] In step 240, the risk identification big model includes a multimodal big model, which can further introduce textual description information for each target video data to enrich the content for risk identification by the multimodal big model.
[0078] Figure 7 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0079] According to some embodiments, such as Figure 7 As shown, the risk identification big model includes a multimodal big model, and method 200 also includes: Step 710: Based on at least one acquisition location associated with at least one target video data, add a text description to at least one target video data to obtain at least one multimodal data corresponding to at least one target video data, wherein the text description indicates potential security risks at at least one acquisition location associated with at least one target video data; and Step 720: Use a multimodal large model to process at least one multimodal data to obtain at least one initial identification result of the security risk monitoring results.
[0080] By adding textual descriptions of the data collection locations, large models can be better equipped to identify risks, thereby improving the accuracy of the identification results.
[0081] In step 710, the text description information may be, for example, the operating procedures for the corresponding acquisition location (e.g., a safe production workshop), thereby providing text information such as violations to the multimodal large model, assisting the large model in identifying whether corresponding violations occur in the target video data.
[0082] For example, corresponding text description information can be added to all or part of at least one target video data to balance the accuracy of the recognition results and the amount of data processing.
[0083] In step 250, since each initial identification result only indicates whether a security risk exists at the corresponding collection location and the type of security risk, a comprehensive security risk monitoring result for the entire target scenario can be generated based on at least one initial identification result. For example, if the scope of the security risk is small and it is eliminated in time, it can be considered that no security risk exists in the target scenario, and only relevant records are saved without risk warnings.
[0084] According to some embodiments, the security risk monitoring results indicate the likelihood of a security risk in the target scenario and the type of security risk. The method 200 further includes: in response to determining that the likelihood of a security risk in the target scenario is higher than a first risk threshold based on the security risk monitoring results, generating a prompt message based on the type of security risk and providing a prompt.
[0085] Based on this, relevant personnel can be promptly alerted to risks, and corresponding safety risks can be addressed to improve the effectiveness of safety production management.
[0086] For example, the first risk threshold could be, for instance, 95%, 98%, or 99%, to provide a security risk warning when there is a high degree of certainty. For example, the warning information could include, for instance, the type of security risk, its location, and the time of occurrence.
[0087] Figure 8 A partial flowchart of another video data processing method according to an embodiment of the present disclosure is shown.
[0088] According to some embodiments, such as Figure 8 As shown, method 200 also includes: Step 810: In response to determining, based on the security risk monitoring results, that the probability of a security risk in the target scenario is lower than a first risk threshold and higher than a second risk threshold, determine at least one collection location associated with the security risk type, wherein the second risk threshold is lower than the first risk threshold; Step 820: Obtain at least one monitoring video data point corresponding to at least one acquisition location associated with a security risk type from multiple monitoring video data points; and Step 830: Process at least one piece of surveillance video data using the risk identification big data model to update the security risk monitoring results.
[0089] In some cases, it may be difficult to accurately determine the security risk situation and type of security risk at the corresponding collection location based solely on the merged target video data. In such cases, the complete surveillance video data can be retrieved for verification, thereby improving the accuracy of security risk monitoring results.
[0090] In step 810, the second risk threshold can be, for example, 50%, 70%, and 80%, so that when it is determined that there is a potential security risk but the certainty is not high, the complete surveillance video data can be retrieved for verification.
[0091] In step 820, multiple surveillance video data can be stored, for example, through a distributed storage system based on minio as described above, thereby enabling the accurate location and retrieval of relevant data from massive amounts of data.
[0092] Figure 9 A flowchart of a model training method according to an embodiment of the present disclosure is shown.
[0093] like Figure 9 As shown, model training method 900 includes: Step 910: Obtain the training dataset, which includes multiple training video data sets and annotation results for the multiple training video data sets. The annotation results indicate the security risk type of the acquisition location associated with the multiple training video data sets; and Step 920: Train the initial large model using the training dataset to obtain the risk identification large model. The risk identification large model is used to identify whether there are security risks associated with the acquisition location of the target video data and the types of security risks that exist.
[0094] Therefore, a large-scale risk identification model can be trained to identify security risks in target video data, thereby realizing the aforementioned video data processing method.
[0095] According to another aspect of this disclosure, a video data processing apparatus is provided. For example... Figure 10As shown, the video data processing device 1000 includes: a first module 1010 configured to acquire multiple monitoring video data corresponding to multiple acquisition locations of a target scene; a second module 1020 configured to extract video segments associated with key events from the multiple monitoring video data to obtain multiple video segments, wherein the key events indicate that the acquisition location of the monitoring video data may have security risks; a third module 1030 configured to merge the multiple video segments to obtain at least one target video data; a fourth module 1040 configured to process at least one target video data using a risk identification big data model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether there is a security risk at the acquisition location associated with the corresponding target video data and the type of security risk present; and a fifth module 1050 configured to determine a security risk monitoring result for the target scene based on at least one initial identification result.
[0096] According to another aspect of this disclosure, a model training apparatus is provided. For example... Figure 11 As shown, the model training device 1100 includes: a sixth module 1110 configured to acquire a training dataset, wherein the training dataset includes multiple training video data and annotation results for the multiple training video data, the annotation results indicating the security risk type of the acquisition location associated with the multiple training video data; and a seventh module 1120 configured to train an initial large model using the training dataset to obtain a risk identification large model, wherein the risk identification large model is used to identify whether there is a security risk associated with the acquisition location of the target video data and the type of security risk.
[0097] According to another aspect of this disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the aforementioned method.
[0098] According to another aspect of this disclosure, a non-transitory computer-readable storage medium storing computer instructions is also provided, wherein the computer instructions are used to cause the computer to perform the aforementioned method.
[0099] According to another aspect of this disclosure, a computer program product is also provided, including a computer program, wherein the computer program implements the aforementioned method when executed by a processor.
[0100] refer to Figure 12The present invention describes a structural block diagram of an electronic device 1200 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0101] like Figure 12 As shown, the electronic device 1200 includes a computing unit 1201, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1202 or a computer program loaded from a storage unit 1208 into a random access memory (RAM) 1203. The RAM 1203 may also store various programs and data required for the operation of the electronic device 1200. The computing unit 1201, ROM 1202, and RAM 1203 are interconnected via a bus 1204. An input / output (I / O) interface 1205 is also connected to the bus 1204.
[0102] Multiple components in electronic device 1200 are connected to I / O interface 1205, including: input unit 1206, output unit 1207, storage unit 1208, and communication unit 1209. Input unit 1206 can be any type of device capable of inputting information to electronic device 1200. Input unit 1206 can receive input digital or character information and generate key signal input related to user settings and / or function control of electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 1207 can be any type of device capable of presenting information, and may include, but is not limited to, a monitor, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 1208 may include, but is not limited to, disk and optical disk. Communication unit 1209 allows electronic device 1200 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth. TM Devices, 802.11 devices, WiFi devices, WiMax devices, cellular communication devices and / or the like.
[0103] The computing unit 1201 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 performs the various methods and processes described above, such as GPU-based matrix computation methods. For example, in some embodiments, the video data processing method can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1208. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 1200 via ROM 1202 and / or communication unit 1209. When the computer program is loaded into RAM 1203 and executed by the computing unit 1201, one or more steps of the video data processing method described above can be performed. Alternatively, in other embodiments, the computing unit 1201 may be configured to perform video data processing methods by any other suitable means (e.g., by means of firmware).
[0104] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0105] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0107] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0108] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0109] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact via communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, servers in distributed systems, or servers incorporating blockchain technology.
[0110] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0111] While embodiments or examples of this disclosure have been described with reference to the accompanying drawings, it should be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the invention is not limited by these embodiments or examples, but only by the granted claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by their equivalents. Furthermore, the steps may be performed in a different order than that described in this disclosure. Further, various elements in the embodiments or examples may be combined in various ways. Importantly, as the technology evolves, many elements described herein can be replaced by equivalents that appear after this disclosure.
Claims
1. A method for processing video data, comprising: Acquire multiple surveillance video data for a target scenario; Video clips associated with key events are extracted from the multiple surveillance video data to obtain multiple video clips, wherein the key events indicate that there may be security risks at the location where the surveillance video data was collected; The multiple video segments are merged to obtain at least one target video data. The at least one target video data is processed using a risk identification model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether there is a security risk associated with the acquisition location of the corresponding target video data and the type of security risk present; and Based on the at least one initial identification result, the security risk monitoring results for the target scenario are determined.
2. The method according to claim 1, wherein, The process of merging the multiple video segments to obtain at least one target video data includes: The multiple video segments are classified based on their acquisition information to obtain at least one set of segments, wherein the acquisition information includes at least one of the acquisition location and acquisition time of the multiple video segments; and The merging process is performed on the plurality of video segments based on the at least one segment set to obtain the at least one target video data corresponding to the at least one segment set.
3. The method according to claim 2, wherein, In response to the acquisition information including the acquisition locations of the multiple video segments, the distance between the acquisition locations of any two video segments in the segment set obtained by classification is less than a distance threshold; as well as In response to the acquisition information including the acquisition time of the multiple video segments, the time difference between the acquisition times of any two video segments in the classified segment set is less than a time threshold.
4. The method according to any one of claims 1-3, wherein, The step of extracting video clips associated with key events from the multiple surveillance video data yields multiple video clips, including: Determine multiple processing tasks corresponding to the multiple monitoring video data for extracting video segments associated with the key events from the monitoring video data; Based on the acquisition location of the multiple monitoring video data corresponding to the multiple processing tasks, a task sequence for the multiple monitoring video data in the thread pool is determined; A task pool is established based on the task sequence to process the multiple processing tasks, resulting in the multiple video segments; and The multiple video segments are stored in their respective storage spaces.
5. The method according to claim 4, wherein, The acquisition locations include main viewpoint locations and auxiliary viewpoint locations. The step of establishing a task pool based on the task sequence to process the multiple processing tasks includes: In response to determining that at least one main view task is included among the plurality of processing tasks, at least one corresponding task pool is established for the at least one main view task, wherein the main view task indicates that the acquisition location of the corresponding monitoring video data is the main view location; Based on the at least one task pool, the at least one main-view task is processed in parallel; In response to determining that a completed main-view task exists, an auxiliary-view task associated with the completed main-view task is obtained from the plurality of processing tasks, wherein the auxiliary-view task indicates that the acquisition location of the corresponding monitoring video data is the auxiliary-view location; and The auxiliary perspective task is processed using the task pool corresponding to the completed main perspective task.
6. The method according to any one of claims 1-5, wherein, The risk identification big model includes a multimodal big model, and the method further includes: Based on at least one acquisition location associated with the at least one target video data, a text description is added to the at least one target video data to obtain at least one multimodal data corresponding to the at least one target video data, wherein the text description indicates potential security risks at the at least one acquisition location associated with the at least one target video data; and The at least one multimodal data is processed using the multimodal large model to obtain the at least one initial recognition result.
7. The method according to any one of claims 1-6, wherein, The security risk monitoring results indicate the likelihood of a security risk in the target scenario and the type of security risk present. The method further includes: In response to determining, based on the security risk monitoring results, that the probability of a security risk in the target scenario is higher than a first risk threshold, a prompt message is generated and issued according to the type of security risk.
8. The method according to claim 7, further comprising: In response to determining, based on the security risk monitoring results, that the probability of a security risk in the target scenario is lower than the first risk threshold and higher than the second risk threshold, at least one collection location associated with the security risk type is determined, wherein the second risk threshold is lower than the first risk threshold; Obtain at least one monitoring video data point corresponding to at least one acquisition location associated with the security risk type from the plurality of monitoring video data points; and The risk identification big data model is used to process the at least one surveillance video data to update the security risk monitoring results.
9. The method according to any one of claims 1-8, wherein, The step of extracting video clips associated with key events from the multiple surveillance video data yields multiple video clips, including: Based on the corresponding acquisition locations of the multiple surveillance video data, an event recognition model is determined, wherein the event recognition model is used to identify the key events; and The event recognition model is used to process the multiple surveillance video data to obtain the multiple video clips.
10. The method according to any one of claims 1-9, wherein, The monitored video data is video stream data.
11. A model training method, comprising: Obtain a training dataset, which includes multiple training video data and annotation results for the multiple training video data, wherein the annotation results indicate the security risk type of the acquisition location associated with the multiple training video data; as well as The initial large model is trained using the training dataset to obtain a risk identification large model, wherein the risk identification large model is used to identify whether there are security risks at the acquisition location associated with the target video data and the type of security risks that exist.
12. A video data processing apparatus, comprising: The first module is configured to acquire multiple surveillance video data corresponding to multiple acquisition locations in the target scene; The second module is configured to extract video segments associated with key events from the multiple surveillance video data to obtain multiple video segments, wherein the key events indicate that the collection location of the surveillance video data may have security risks; The third module is configured to merge the multiple video segments to obtain at least one target video data. The fourth module is configured to process the at least one target video data using a risk identification model to obtain at least one initial identification result, wherein the at least one initial identification result indicates whether a security risk exists at the acquisition location associated with the corresponding target video data and the type of security risk present; and The fifth module is configured to determine the security risk monitoring results for the target scenario based on the at least one initial identification result.
13. A model training device, comprising: The sixth module is configured to acquire a training dataset, wherein the training dataset includes multiple training video data sets and annotation results for the multiple training video data sets, the annotation results indicating the security risk type of the acquisition location associated with the multiple training video data sets; and The seventh module is configured to train an initial large model using the training dataset to obtain a risk identification large model, wherein the risk identification large model is used to identify whether there is a security risk at the acquisition location associated with the target video data and the type of security risk.
14. An electronic device comprising: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.
15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.
16. A computer program product comprising a computer program, wherein, When the computer program is executed by a processor, it implements the method of any one of claims 1-11.