Data acquisition processing evaluation method and system based on video image
By collecting and detecting video data in real time, using the preset video analysis model to generate image analysis data, and sorting it into text interpretation results through text interpretation templates, the problems of low efficiency, delay time and result deviation of manual analysis of video images are solved, and efficient and accurate video data analysis is achieved.
Patent Information
- Application Number
- CN202510187233.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, manual analysis of video images leads to high labor intensity, low efficiency, long delay, and susceptible to personal subjective factors, resulting in large deviations in the results and insufficient intelligence on-site applications.
Video data is collected in real time through the video acquisition device, and the video data is detected according to the preset time interval to determine each target object in the video data. The preset video analysis model is used to analyze each image frame marked with the target object video, generate image analysis data, and organize it into text interpretation results through text interpretation templates, and display it on the display screen of the display console.
It greatly improves the speed and efficiency of video data analysis, improves the accuracy of video analysis, avoids the influence of human subjective factors, ensures the objectivity of image analysis data, and solves the problems of high labor intensity, low efficiency, long delay and result deviation of manual analysis of video images.
Smart Images

Figure CN120034627A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of video image processing, and in particular to a method and system for evaluating data acquisition, processing and evaluation based on video images. Background Art
[0002] High-definition video has been widely used in daily monitoring. These advanced monitoring methods capture video images from different angles, making the monitoring data more complete. To a certain extent, they make up for the shortcomings of manual monitoring and improve work efficiency and benefits. However, the large number of monitoring video images generated lack automated and intelligent data processing methods.
[0003] At present, for the videos collected by the video acquisition equipment, the operator needs to operate or judge according to the screen display, or perform post-analysis according to the recorded message data. For operators who specialize in video analysis, to a certain extent, they can directly obtain the required data through the video analysis collected by the video acquisition equipment, but for operators who are not specialized in video analysis, they need to replay the screen display, and replay and record frame by frame to obtain the required data. The manual video analysis method is not only labor-intensive, inefficient, and has a long delay, but also the manual analysis of video images is easily affected by personal subjective factors, resulting in large deviations in the final results and insufficient intelligence in on-site applications. Summary of the invention
[0004] The purpose of the present invention is to provide a method and system for data acquisition, processing and evaluation based on video images, so as to improve the problems of high labor intensity, low efficiency and long delay caused by manual video analysis in the prior art.
[0005] The embodiment of the present invention is achieved as follows:
[0006] In a first aspect, an embodiment of the present application provides a method for data acquisition, processing and evaluation based on video images, comprising the following steps:
[0007] The video data is collected in real time by a video collection device, and the video data is detected at a preset time interval to determine each target object in the video data; wherein the video data includes a plurality of image frames, and each image frame includes corresponding collection time information;
[0008] After marking each target object contained in each image frame, based on the acquisition time information of each image frame, each image frame marked with the target object is sequentially input into a preset video analysis model for video analysis to obtain image analysis data corresponding to each image frame; wherein the image analysis data corresponding to each image frame includes a plurality of image element data;
[0009] According to a predetermined text interpretation template, the image element data corresponding to each image frame is sorted to obtain the text interpretation result corresponding to each image frame;
[0010] Based on the text interpretation results corresponding to each image frame, the image analysis data of each image frame is displayed on the display screen of the display control console; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area being used to display the text interpretation results corresponding to different image element data;
[0011] Based on the video evaluation standard, a video evaluation instruction is sent to the display and control console, and the video evaluation result input by the user is recovered based on the display and control console.
[0012] In some embodiments of the present invention, the detecting of video data at a preset time interval includes:
[0013] Extract corresponding video data at preset time intervals, and detect the video data using target detection technology to determine whether the video data contains a target object;
[0014] If the video data contains the target object, all the target objects in the video data are obtained.
[0015] In some embodiments of the present invention, the determination rules of the above-mentioned text interpretation template include:
[0016] Input all target objects contained in the video data into the preset template database for matching, and determine the text interpretation sub-template corresponding to each target object;
[0017] The text interpretation sub-templates corresponding to each target object are integrated to obtain the text interpretation template corresponding to the video data.
[0018] In some embodiments of the present invention, before inputting all target objects contained in the video data into the preset template database for matching, the method further includes:
[0019] Obtain the historical video data and corresponding text interpretation information corresponding to each target object;
[0020] Generate a text interpretation sub-template corresponding to each target object according to the historical video data and the corresponding text interpretation information corresponding to each target object;
[0021] The text interpretation sub-templates corresponding to all target objects are integrated to establish a preset template database.
[0022] In some embodiments of the present invention, the above-mentioned image element data at least includes one or more of speed data, orientation data of each target object and distance data between each target object.
[0023] In some embodiments of the present invention, the establishment rules of the preset video parsing model include:
[0024] Establish an initial model for video parsing;
[0025] Acquire multiple samples, the samples including historical image data collected by a video acquisition device;
[0026] The initial video parsing model is trained using multiple samples to obtain a preset video parsing model.
[0027] In some embodiments of the present invention, the above-mentioned video analysis initial model is a random forest model or a neural network model.
[0028] In a second aspect, an embodiment of the present application provides a data acquisition, processing and evaluation system based on video images, comprising:
[0029] A target object determination module is used to collect video data in real time through a video acquisition device, and detect the video data at a preset time interval to determine each target object in the video data; wherein the video data includes multiple image frames, and each image frame includes corresponding acquisition time information;
[0030] The video parsing module is used to mark each target object contained in each image frame, and then input each image frame marked with the target object into a preset video parsing model for video parsing based on the acquisition time information of each image frame, so as to obtain image parsing data corresponding to each image frame; wherein the image parsing data corresponding to each image frame includes multiple image element data;
[0031] A text interpretation module, used to sort out the image element data corresponding to each image frame according to a predetermined text interpretation template, and obtain the text interpretation result corresponding to each image frame;
[0032] An image analysis data display module is used to display the image analysis data of each image frame on the display screen of the display control console based on the text interpretation results corresponding to each image frame; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area being used to display the text interpretation results corresponding to different image element data;
[0033] The video evaluation result recovery module is used to send a video evaluation instruction to the display and control console based on the video evaluation standard, and to recover the video evaluation result input by the user based on the display and control console.
[0034] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory for storing one or more programs and a processor. When the one or more programs are executed by the processor, any method in the first aspect is implemented.
[0035] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method as described in any one of the first aspects above.
[0036] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:
[0037] The present invention provides a data acquisition, processing and evaluation method and system based on video images. The video data is collected in real time by a video acquisition device, and the video data is detected at a preset time interval to determine each target object in the video data. After marking each target object contained in each image frame of the video data, each image frame marked with the target object is analyzed in turn by using a preset video analysis model, so as to obtain a variety of image element data of each image frame. Then, the image analysis data corresponding to each image frame is automatically generated through image recognition, which avoids the user from replaying and recording the video data frame by frame, greatly improves the speed and efficiency of analyzing the video data, and performs video analysis by using a preset video analysis model, which can effectively improve the accuracy of video analysis, avoid the influence of human subjective factors, and ensure the objectivity of the image analysis data corresponding to each image frame. According to a predetermined text interpretation template, the image element data corresponding to each image frame is sorted to obtain the text interpretation result corresponding to each image frame. For each image frame, the text interpretation result corresponding to the different image element data of each image frame is presented in the corresponding display area of the display screen of the display console to intuitively show the information contained in the video data to the user. Finally, the video evaluation results input by the user are collected based on the display console. This effectively solves the problems of high labor intensity, low efficiency, long delay caused by manual video analysis, and the fact that manual video image analysis is easily affected by personal subjective factors, resulting in large deviations in the final results and insufficient intelligence in on-site applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 A flowchart of a method for data acquisition, processing and evaluation based on video images provided by an embodiment of the present invention;
[0040] Figure 2 A block diagram of a data acquisition, processing and evaluation system based on video images provided by an embodiment of the present invention;
[0041] Figure 3 A schematic structural block diagram of an electronic device provided by an embodiment of the present invention.
[0042] Icon: 101 - memory; 102 - processor; 103 - communication interface. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. The components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0044] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for which protection is sought, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in the field without creative work are within the scope of protection of the present application.
[0045] It should be noted that similar reference numerals and letters represent similar items in the following drawings, so once an item is defined in one drawing, it does not need to be further defined and explained in the subsequent drawings. At the same time, in the description of this application, if the terms "first", "second", etc. appear, they are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0046] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, if the terms "include", "comprise" or any other variant thereof appear, it is intended to cover non-exclusive inclusion, so that the process, method, article or equipment including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or equipment. In the absence of further restrictions, if an element defined by the sentence "comprises a..." appears, it does not exclude the presence of other identical elements in the process, method, article or equipment including the element.
[0047] In the description of the present application, it should be noted that if the terms "upper", "lower", "inside", "outside", etc. appear to indicate an orientation or position relationship, they are based on the orientation or position relationship shown in the drawings, or are the orientation or position relationship in which the product of the application is usually placed when used. This is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present application.
[0048] In the description of this application, it is also necessary to explain that, unless otherwise clearly specified and limited, the terms "set" and "connection" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection, or it can be an indirect connection through an intermediate medium, or it can be the internal connection of two components. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0049] In conjunction with the accompanying drawings, some implementation methods of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.
[0050] Example
[0051] Please refer to Figure 1 , Figure 1 The present invention provides a flow chart of a data acquisition, processing and evaluation method based on video images. The present invention provides a data acquisition, processing and evaluation method based on video images, including the following steps:
[0052] S110: collecting video data in real time through a video acquisition device, and detecting the video data at a preset time interval to determine each target object in the video data; wherein the video data includes a plurality of image frames, and each image frame includes corresponding acquisition time information;
[0053] In some implementations of the present embodiment, the above-mentioned detection of video data at preset time intervals includes: extracting corresponding video data at preset time intervals, detecting the video data using target detection technology to determine whether the video data contains a target object; if the video data contains a target object, obtaining all target objects in the video data.
[0054] In some implementations of this embodiment, the method further includes: presetting a plurality of target objects according to user requirement parameters. Therefore, when detecting the video data, only the pre-set target objects are detected in the video data.
[0055] Specifically, in the process of real-time video data acquisition by the video acquisition device, the video data is segmented and extracted according to preset time intervals, and the target detection technology is used to detect the target objects in the segmented extracted video data to more accurately locate the target objects, thereby obtaining all target objects contained in the video data.
[0056] S120: After marking each target object contained in each image frame, based on the acquisition time information of each image frame, the image frames marked with the target object are sequentially input into a preset video analysis model for video analysis to obtain image analysis data corresponding to each image frame; wherein the image analysis data corresponding to each image frame includes a plurality of image element data;
[0057] In some implementations of this embodiment, the above-mentioned image element data at least includes one or more of speed data, orientation data of each target object and distance data between each target object.
[0058] Specifically, after marking each target object contained in each image frame, the preset video analysis model is used to analyze each image frame marked with the target object in turn, thereby obtaining various image element data such as speed data, azimuth data and distance data between each target object. This also realizes the recognition of image elements in video data (for example, time, moment, speed data, azimuth data, emission information and distance data between each target object, etc.) and then automatically generates image analysis data corresponding to each image frame through image recognition, avoiding the user from replaying and recording the video data frame by frame, greatly improving the speed and efficiency of analyzing the video data, and performing video analysis through the preset video analysis model can effectively improve the accuracy of video analysis, avoid the influence of human subjective factors, and ensure the objectivity of the image analysis data corresponding to each image frame.
[0059] S130: sorting the image element data corresponding to each image frame according to a predetermined text interpretation template to obtain a text interpretation result corresponding to each image frame;
[0060] In some implementations of the present embodiment, the rules for determining the above-mentioned text interpretation template include: inputting all target objects contained in the video data into a preset template database for matching, and determining the text interpretation sub-template corresponding to each target object; and integrating the text interpretation sub-templates corresponding to each target object to obtain the text interpretation template corresponding to the video data.
[0061] The preset template database contains text interpretation sub-templates corresponding to each target object.
[0062] Specifically, by typeset and sorting the image element data corresponding to each image frame according to the text interpretation template determined according to each target object contained in the video data, the text interpretation result corresponding to each image frame can be more matched with the target object contained in the video data.
[0063] S140: Based on the text interpretation results corresponding to each image frame, the image analysis data of each image frame is displayed on the display screen of the display control console; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area being used to display the text interpretation results corresponding to different image element data;
[0064] Specifically, the display screen of the display console is divided into multiple display areas, and each display area displays different image element data (for example, display area A is used to display target object speed data, azimuth data, and launch information, etc., display area B is used to display equipment status information, and display area C is used to display weapon strike information). For each image frame, the text interpretation results corresponding to the different image element data of each image frame are presented in the corresponding display area of the display screen of the display console to intuitively show the information contained in the video data to the user.
[0065] S150: Based on the video evaluation standard, a video evaluation instruction is sent to the display control console, and the display control console recovers the video evaluation result input by the user.
[0066] Specifically, the method collects video data in real time through a video acquisition device, and detects the video data at a preset time interval to determine each target object in the video data. After marking each target object contained in each image frame of the video data, the preset video analysis model is used to analyze each image frame marked with the target object in turn, thereby obtaining a variety of image element data of each image frame. Then, the image analysis data corresponding to each image frame is automatically generated through image recognition, which avoids the user from replaying and recording the video data frame by frame, greatly improving the speed and efficiency of analyzing the video data, and performing video analysis through the preset video analysis model can effectively improve the accuracy of video analysis, avoid the influence of human subjective factors, and ensure the objectivity of the image analysis data corresponding to each image frame. According to the predetermined text interpretation template, the image element data corresponding to each image frame is sorted to obtain the text interpretation result corresponding to each image frame. For each image frame, the text interpretation result corresponding to the different image element data of each image frame is presented in the corresponding display area of the display screen of the display console to intuitively show the information contained in the video data to the user. Finally, the video evaluation result input by the user is recovered based on the display console. It effectively solves the problems of high labor intensity, low efficiency and long delay caused by manual video analysis, and the fact that manual video image analysis is easily affected by personal subjective factors, resulting in large deviations in the final results and insufficient intelligence in on-site applications.
[0067] In some implementations of the present embodiment, before inputting all target objects contained in the video data into the preset template database for matching, the method also includes: obtaining historical video data and corresponding text interpretation information corresponding to each target object; generating a text interpretation sub-template corresponding to each target object based on the historical video data and corresponding text interpretation information corresponding to each target object; and establishing a preset template database by integrating the text interpretation sub-templates corresponding to all target objects.
[0068] Specifically, the text interpretation sub-templates corresponding to each target object are summarized and sorted according to the historical video data and the corresponding text interpretation information corresponding to each target object. Thus, the text interpretation sub-templates corresponding to all target objects are integrated to establish a preset template database, so that the preset template database contains the text interpretation sub-templates corresponding to each target object.
[0069] In some implementations of the present embodiment, the establishment rules of the above-mentioned preset video analysis model include: establishing an initial video analysis model; obtaining multiple samples, the samples include historical image data collected by a video acquisition device; using multiple samples to train the initial video analysis model to obtain a preset video analysis model.
[0070] In some implementations of this embodiment, the above-mentioned video analysis initial model is a random forest model or a neural network model.
[0071] Please refer to Figure 2 , Figure 2 A block diagram of a data acquisition, processing and evaluation system based on video images provided by an embodiment of the present invention. The present application embodiment provides a data acquisition, processing and evaluation system based on video images, including:
[0072] A target object determination module is used to collect video data in real time through a video acquisition device, and detect the video data at a preset time interval to determine each target object in the video data; wherein the video data includes multiple image frames, and each image frame includes corresponding acquisition time information;
[0073] The video parsing module is used to mark each target object contained in each image frame, and then input each image frame marked with the target object into a preset video parsing model for video parsing based on the acquisition time information of each image frame, so as to obtain image parsing data corresponding to each image frame; wherein the image parsing data corresponding to each image frame includes multiple image element data;
[0074] A text interpretation module, used to sort out the image element data corresponding to each image frame according to a predetermined text interpretation template, and obtain the text interpretation result corresponding to each image frame;
[0075] An image analysis data display module is used to display the image analysis data of each image frame on the display screen of the display control console based on the text interpretation results corresponding to each image frame; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area being used to display the text interpretation results corresponding to different image element data;
[0076] The video evaluation result recovery module is used to send a video evaluation instruction to the display and control console based on the video evaluation standard, and to recover the video evaluation result input by the user based on the display and control console.
[0077] Specifically, the system collects video data in real time through a video acquisition device, and detects the video data at a preset time interval to determine each target object in the video data. After marking each target object contained in each image frame of the video data, the preset video analysis model is used to analyze each image frame marked with the target object in turn, thereby obtaining a variety of image element data of each image frame. Then, the image analysis data corresponding to each image frame is automatically generated through image recognition, which avoids the user from replaying and recording the video data frame by frame, greatly improving the speed and efficiency of analyzing the video data, and performing video analysis through the preset video analysis model can effectively improve the accuracy of video analysis, avoid the influence of human subjective factors, and ensure the objectivity of the image analysis data corresponding to each image frame. According to the predetermined text interpretation template, the image element data corresponding to each image frame is sorted to obtain the text interpretation results corresponding to each image frame. For each image frame, the text interpretation results corresponding to the different image element data of each image frame are presented in the corresponding display area of the display screen of the display console to intuitively show the information contained in the video data to the user. Finally, the video evaluation results input by the user are recovered based on the display console. It effectively solves the problems of high labor intensity, low efficiency and long delay caused by manual video analysis, and the fact that manual video image analysis is easily affected by personal subjective factors, resulting in large deviations in the final results and insufficient intelligence in on-site applications.
[0078] Please refer to Figure 3 , Figure 3 A schematic structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device includes a memory 101, a processor 102 and a communication interface 103, and the memory 101, the processor 102 and the communication interface 103 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, these components can be electrically connected to each other through one or more communication buses or signal lines. The memory 101 can be used to store software programs and modules, such as program instructions / modules corresponding to a data acquisition, processing and evaluation system based on video images provided in an embodiment of the present application, and the processor 102 executes various functional applications and data processing by executing software programs and modules stored in the memory 101. The communication interface 103 can be used for signaling or data communication with other node devices.
[0079] Among them, the memory 101 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable read-only memory (EEPROM), etc.
[0080] The processor 102 may be an integrated circuit chip with signal processing capability. The processor 102 may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0081] Understandably, Figure 3 The structure shown is for illustration only. The electronic device may also include Figure 3 More or fewer components as shown, or with Figure 3 Different configurations shown. Figure 3 Each component shown in the figure can be implemented by hardware, software or a combination thereof.
[0082] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, a program segment or a part of a code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of boxes in the block diagram and / or the flowchart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0083] In addition, the functional modules in the various embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0084] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0085] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0086] It will be apparent to those skilled in the art that the present application is not limited to the details of the exemplary embodiments described above, and that the present application can be implemented in other specific forms without departing from the spirit or essential features of the present application. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the present application is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present application. Any reference numeral in a claim should not be considered as limiting the claim to which it relates.
Claims
1. A data acquisition, processing and evaluation method based on video images, characterized in that: The steps include: The video data is collected in real time by a video collection device, and the video data is detected at a preset time interval to determine each target object in the video data; wherein the video data includes a plurality of image frames, and each image frame includes corresponding collection time information; After marking each target object contained in each image frame, based on the acquisition time information of each image frame, each image frame marked with the target object is sequentially input into a preset video analysis model for video analysis to obtain image analysis data corresponding to each image frame; wherein the image analysis data corresponding to each image frame includes a plurality of image element data; According to a predetermined text interpretation template, the image element data corresponding to each image frame is sorted to obtain the text interpretation result corresponding to each image frame; Based on the text interpretation results corresponding to each image frame, the image analysis data of each image frame is displayed on the display screen of the display control console; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area being used to display the text interpretation results corresponding to different image element data; Based on the video evaluation standard, a video evaluation instruction is sent to the display and control console, and the video evaluation result input by the user is recovered based on the display and control console.
2. The video image-based data acquisition, processing and evaluation method according to claim 1, characterized in that: The detecting the video data at a preset time interval includes: Extracting corresponding video data at a preset time interval, and detecting the video data using a target detection technology to determine whether the video data contains a target object; If the video data contains target objects, all target objects in the video data are acquired.
3. The video image-based data acquisition, processing and evaluation method according to claim 1, characterized in that: The determination rules of the text interpretation template include: Input all target objects contained in the video data into the preset template database for matching, and determine the text interpretation sub-template corresponding to each target object; The text interpretation sub-templates corresponding to each target object are integrated to obtain the text interpretation template corresponding to the video data.
4. The video image-based data acquisition, processing and evaluation method according to claim 3, characterized in that: Before inputting all target objects contained in the video data into the preset template database for matching, the method further includes: Obtain the historical video data and corresponding text interpretation information corresponding to each target object; Generate a text interpretation sub-template corresponding to each target object according to the historical video data and the corresponding text interpretation information corresponding to each target object; The text interpretation sub-templates corresponding to all target objects are integrated to establish a preset template database.
5. The video image-based data acquisition, processing and evaluation method according to claim 1, characterized in that: The image element data at least includes one or more of speed data, orientation data of each target object and distance data between each target object.
6. The video image-based data acquisition, processing and evaluation method according to claim 1, characterized in that: The establishment rules of the preset video analysis model include: Establish an initial model for video parsing; Acquire a plurality of samples, wherein the samples include historical image data acquired by a video acquisition device; The multiple samples are used to train the initial video parsing model to obtain a preset video parsing model.
7. The video image-based data acquisition, processing and evaluation method according to claim 6, characterized in that: The video analysis initial model is a random forest model or a neural network model.
8. A data acquisition, processing and evaluation system based on video images, characterized in that: include: A target object determination module is used to collect video data in real time through a video acquisition device, and detect the video data at preset time intervals to determine each target object in the video data; wherein the video data includes multiple image frames, and each image frame includes corresponding acquisition time information; The video parsing module is used to mark each target object contained in each image frame, and then input each image frame marked with the target object into a preset video parsing model for video parsing based on the acquisition time information of each image frame, so as to obtain image parsing data corresponding to each image frame; wherein the image parsing data corresponding to each image frame includes multiple image element data; A text interpretation module, used to sort out the image element data corresponding to each image frame according to a predetermined text interpretation template, and obtain the text interpretation result corresponding to each image frame; An image analysis data display module is used to display the image analysis data of each image frame on the display screen of the display control console based on the text interpretation results corresponding to each image frame; wherein the display control console is communicatively connected with the video acquisition device, and the display screen of the display control console is divided into a plurality of display areas, each display area is used to display the text interpretation results corresponding to different image element data; The video evaluation result recovery module is used to send a video evaluation instruction to the display and control console based on the video evaluation standard, and to recover the video evaluation result input by the user based on the display and control console.
9. An electronic device, characterized in that: include: A memory for storing one or more programs; processor; When the one or more programs are executed by the processor, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.