A method, system, device, server, and medium for acquiring training data.
By processing video data and automatically labeling, the problem of motion category detection being difficult to adapt to rapid changes in existing technologies has been solved, resulting in improved accuracy of motion classification models and recognition of new motion categories, while reducing training costs.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-28
- Publication Date
- 2026-03-06
AI Technical Summary
Existing action category detection technologies struggle to quickly increase the number of action categories detected or improve detection accuracy in rapidly changing application scenarios, and they cannot control the leakage of information through photos in real time.
By acquiring video data, object detection models and action classification models are used to classify actions. Image frames that are identified as erroneous are stored and automatically labeled. Category labels are configured to achieve rapid training data acquisition and improve the accuracy of action classification models.
It enables rapid training data acquisition for action classification models, improves the detection accuracy of existing action categories, and can identify new action categories, thereby reducing training costs.
Smart Images

Figure CN117292222B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data acquisition technology, and specifically to a method, system, device, server, and medium for acquiring training data. Background Technology
[0002] In today's corporate office environment, employees complete most of their work via computers and networks. This results in a large amount of confidential data being stored as electronic files on employee computers, and many trade secrets and intangible assets require computer storage. Preventing these files from being disseminated and avoiding incalculable losses to the company is a top priority for businesses in preventing data leaks. Although most companies have adopted many measures to prevent data leaks, such as establishing a pure intranet office environment, prohibiting the use of mobile storage, monitoring employee computer operations, prohibiting cloud storage sharing, and controlling file sharing, they have all overlooked a crucial leak route: photographic leaks. While adding watermarks to application windows can facilitate leak tracing, it cannot control leaks in real time or immediately. Therefore, in such cases, when a leaker takes a photograph, it should be immediately identified for timely behavioral control.
[0003] In existing technologies, when action category detection is required, existing action category detection schemes can be used to implement spatio-temporal action detection tasks. However, this method has a problem: in practical applications, the action categories to be detected are often not static; there is a tendency to add more action categories or further improve the detection accuracy of existing action categories. When adding action categories or further improving the detection accuracy of existing action categories, it is often necessary to collect and label large amounts of data. This method is difficult to adapt to application scenarios with rapidly changing requirements. Summary of the Invention
[0004] This invention provides a method, system, device, server, and medium for collecting training data, which can effectively improve the detection accuracy of existing action categories and increase the number of action categories to be detected.
[0005] In a first aspect, the present invention provides a method for collecting training data, comprising:
[0006] Acquire video data of the test actions;
[0007] The video data of the test action is input into a trained object detection model to perform object detection, and the object detection result is obtained; the object detection result is input into a trained action classification model to perform action classification, and the action classification result is obtained.
[0008] When the action classification result identifies an action image that does not belong to the preset storage area, the action image is stored in the preset storage area and automatically labeled. The storage area is configured with category identifiers for existing actions, and the category identifiers correspond to the action category of the test action being performed.
[0009] The above technical solution acquires video data, classifies the video data into actions based on an action classification model, and collects image frames that are incorrectly classified by the action classification model. These frames are then stored in a preset storage area and automatically labeled, thereby enabling rapid acquisition of training data for the action classification model and improving the detection accuracy of existing action categories.
[0010] Optionally, the data acquisition method further includes:
[0011] Acquire video data of the same new action being performed, wherein the new action is an action of an action category that is not present in the training data of the action classification model;
[0012] The video data of the new action is stored in the storage area corresponding to the new action and automatically labeled. The storage area is equipped with a category identifier for the new action.
[0013] Video data and corresponding category identifiers are extracted from the storage area to train an action classification model, thereby recognizing new actions.
[0014] Alternatively, the same existing action can be performed continuously during the data collection process.
[0015] Optionally, the object detection model uses the object detection model of the YOLO series system.
[0016] Optionally, the action classification model uses a convolutional neural network, a recurrent neural network, or a spatio-Temporal Attention Network model.
[0017] Optionally, the data acquisition method further includes: real-time display and confidence assessment of the actually performed actions and the identified actions.
[0018] Secondly, the present invention also provides a training data acquisition device, comprising:
[0019] The acquisition module is used to acquire video data of the test actions;
[0020] The action detection and recognition module is used to input the video data of the test action into a trained object detection model to perform object detection and obtain object detection results; and to input the object detection results into a trained action classification model to perform action classification and obtain action classification results.
[0021] The storage and annotation module is used to store the action image in the preset storage area and automatically annotate it when the action classification result identifies an action image that does not belong to the preset storage area. The storage area is configured with category identifiers for existing actions, and the category identifiers correspond to the action category of the test action being performed.
[0022] Thirdly, the present invention also provides a training data acquisition system, comprising:
[0023] Video capture devices are used to continuously capture the actions being performed and upload the captured data;
[0024] The server is used to receive the collected data and identify the action category. The server is equipped with a pre-trained object detection model and an action classification model.
[0025] The client provides a real-time recognition interface to display the actions performed and the recognition results given by the model in real time; and provides a confidence score area to display the confidence value in real time.
[0026] Fourthly, the present invention also provides a server, including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the training data acquisition method.
[0027] Fifthly, the present invention also provides a computer-readable storage medium storing a computer program thereon, wherein when the computer program is executed by a processor, it implements the steps of the training data acquisition method.
[0028] Compared with the prior art, the beneficial effects of the present invention are:
[0029] 1. This invention acquires video data, classifies the video data into actions based on an action classification model, and collects image frames that are incorrectly classified by the action classification model. These frames are then stored in a preset storage area and automatically labeled, thereby enabling rapid acquisition of training data for the action classification model and improving the detection accuracy of existing action categories.
[0030] 2. By acquiring video data of new actions, configuring category labels for new actions in the storage area, storing video data and performing automatic annotation, and extracting video data and category labels, the action classification model is trained to add new action categories.
[0031] 3. This invention only trains the action classification model, and only adds training data related to action categories to train the action classification model, thereby reducing training costs. Attached Figure Description
[0032] Figure 1 This is a flowchart illustrating a training data acquisition method according to the present invention;
[0033] Figure 2 This is a schematic diagram illustrating the process of adding and identifying new action categories during the training data acquisition process of this invention.
[0034] Figure 3 This is a schematic diagram illustrating an example of the action classification task based on video data according to the present invention;
[0035] Figure 4 This is a schematic diagram illustrating an example of a training data acquisition method according to the present invention. Detailed Implementation
[0036] It should be noted that in many situations, it is necessary to identify the actions of the actors. For example, in unmanned examination rooms, it is necessary to identify whether the examinee's actions are in violation of regulations, such as looking at cheat sheets or checking their mobile phone. In some special office scenarios, such as the service halls of banks and other financial institutions, it is necessary to identify whether on-site service personnel are eating or sleeping, as these actions can damage the company's image. The training data collection method described in this invention is applied to the above scenarios. For sensitive or prohibited actions, it is necessary to improve the accuracy of action classification, quickly analyze and process the actions performed by the target, and achieve action classification, thereby improving the management efficiency of sensitive or prohibited actions.
[0037] Please explain the following terms:
[0038] Existing actions: The storage area already stores actions corresponding to the action category, that is, actions that action classification model 8 can recognize.
[0039] New action: Actions that do not have an action category in the training data of action classification model 8.
[0040] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.
[0041] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0042] This invention provides a method for collecting training data, comprising:
[0043] Step 1: Obtain video data of the existing motion.
[0044] In this embodiment, a video capture device 12 can be set up. The test subject 10 continuously performs existing actions within the shooting area of the video capture device 12, such as playing with a mobile phone, copying, eating, or other actions. After the video capture device 12 captures the video data, it transmits the video data to the server 11 via the network. It should be noted that there can be multiple types of existing actions, but the test subject 10 can only perform one action in a single capture. In this embodiment, the existing action is one that the action classification model 8 can recognize.
[0045] Step 2: Input the video data of the existing action into the trained object detection model 7 to perform object detection and obtain the object detection result; input the object detection result into the trained action classification model 8 to perform action classification and obtain the action classification result.
[0046] In this embodiment, the two models implement two-stage detection: Stage 1: Human body detection stage: the area is detected by setting an object recognition model, which is to draw a frame around the person; Stage 2: Action classification and recognition stage: the action classification model is set to detect the type of human action.
[0047] In this embodiment, after obtaining the video data, the server 11 sequentially inputs the video data into the pre-trained object detection model 7 and action classification model 8 to perform object detection and action classification.
[0048] In this embodiment, the object detection model 7 can be a YOLO system series object detection model, such as the YOLOv5 model. This YOLOv5 model can be obtained using training data, such as images of people, specifically images of people indoors, such as an office. In this embodiment, the preferred training data is images of people sitting in an office, facing the video capture device 12. The corresponding annotations are for people, that is, drawing frames around the people in the images.
[0049] In this embodiment, the action classification model 8 can employ a convolutional neural network or a recurrent neural network, or a spatio-temporal attention network model. Correspondingly, the training data for the action classification model 8 can be obtained using training data such as images of people, labeled with the action category of the person in the image. In this embodiment, the training data for the action classification model 8 is preferably images of people in an office, with the person facing the video capture device 12 from their perspective. The corresponding label is the action category of the person in the image, which corresponds to a specific scenario. Specifically, the action category is a sensitive or prohibited action in a certain scenario. For example, in a scenario to prevent data leakage, the action category could be the action of a person copying text, or the action of a person taking a picture of a computer screen, etc. These actions could lead to data leakage and are therefore sensitive actions. In a scenario to prevent cheating on an exam, the action category could also be the action of playing on a mobile phone. In some scenarios, the action category could also be the action of eating. For example, in service positions of some large organizations, it is necessary to monitor the actions of employees in real time to prevent employees from performing prohibited actions that could affect the organization's image.
[0050] Step 3: When the action classification result identifies an action image that does not belong to the preset storage area, the action image is stored in the preset storage area and automatically labeled. The storage area is configured with the category identifier of the existing action, and the category identifier corresponds to the action category of the existing action being performed.
[0051] In this embodiment, an identification error can be represented as the action category that should be identified as the existing action in step S1, but the action classification model 8 fails to identify it as such. For example, the action classification model 8 can identify the three action categories A, B, and C. Correspondingly, in step S1, the test subject 10 performs action A, but the model outputs B or C, or it cannot give a specific category. That is, the action classification model 8 believes that the actions performed by the test subject 10 in step S1 do not belong to the three action categories A, B, and C.
[0052] In this embodiment, the storage area corresponding to the existing action is pre-configured. The storage area can be a folder, a storage area directly set in the local storage space of server 11, such as the hard drive of local server 11, or a storage area set in a remote repository, such as a file system (FS), a distributed file system (HDFS), a network file system (NFS), or network attached storage (NAS), or it can be stored in cloud storage, such as Amazon S3, Google Cloud Storage, etc.
[0053] The aforementioned storage areas are configured with category identifiers for existing actions. This can be achieved by creating a folder in the various storage spaces mentioned above, with the folder name serving as the category identifier. For example, in step S1, if tester 10 is performing the action of playing on a mobile phone, then the file category name could be "Playing on a Mobile Phone Action".
[0054] In this embodiment, in step S1, the tester 10 continuously performs the same existing action. Correspondingly, the action classification model 8 should output the same action classification result, and this action classification result should correspond to the existing action. When a certain frame of image is not identified as the category of the existing action, it indicates that the action classification model 8 may lack training data for that action, so it needs to be collected. Therefore, in this embodiment, only the image frames that are incorrectly identified are collected as training data (we believe that for the correctly identified image frames, the action classification model 8 has already learned the corresponding data representation, so it is not necessary to continue collecting them, only the incorrectly identified ones are collected). Secondly, since it is known in advance what existing action the tester 10 will perform, and a corresponding storage area carrying the existing action category identifier is preset, it can be directly known that the image frames in this storage area are all of a certain existing action category, and the action classification model 8 has made a mistake in identifying them. In other words, the action classification model 8 has not learned or has not learned sufficiently the data representation of the image frames in this storage area.
[0055] Therefore, this embodiment can quickly collect images that are misclassified or cannot be classified by the action classification model 8 through the above ingenious method, and automatically label the images (placing the images in the storage area is labeling), thereby realizing the rapid collection of training data for the action classification model 8. In practical applications, the images and corresponding category identifiers stored in the above storage area can be used as labels for training data to train the action classification model 8, thereby improving the accuracy of the action classification model 8 in existing action classification.
[0056] Considering that human actions are often continuous—in other words, a person cannot immediately switch from drinking water to using a mobile phone in the next moment—there must be a continuous transition process. For example, the action of drinking water involves several steps: picking up the cup, bringing it to the lips, drinking, and putting the cup down. These steps represent the semantics of the action of drinking water. Therefore, in a time series, if an image frame is misidentified at a certain moment, it can be assumed that the image frames before and after that moment have some similar data representations. Even if the image frames before and after the misidentified image are correctly identified, it can be considered that the model has not fully learned this data representation, and therefore, further learning is needed.
[0057] Therefore, in this situation, for the aforementioned video data, if a frame is misidentified, several image frames can be searched forward or backward within a certain time window, starting from the misidentified image frame, and placed into the aforementioned storage area as training data. This time window can be 1 second or 0.5 seconds. Besides using a time window, the system can also iterate forward and / or backward from the misidentified image frame to obtain the confidence score for each image frame. This confidence score is also output by the action classification model 8. Image frames with confidence scores below a certain threshold are placed into the aforementioned storage area as training data. This threshold can be 75%.
[0058] The data acquisition method also includes:
[0059] In some cases, as user needs change, it may be necessary to add the recognition of new action categories. For example, the current action classification model 8 can recognize three types of actions: A, B, and C. If a user extracts a requirement and wants the model to recognize action D, then for such situations, some implementations of the above-mentioned training data acquisition method further include the following steps:
[0060] Step 4: Obtain video data of the new action performed by test subject 10, wherein test subject 10 continuously performs the same new action. In this embodiment, the new action represents an action of an action category not present in the training data of the action classification model 8.
[0061] Step 5: Store the video data of the new action into the storage area corresponding to the new action to achieve automatic annotation of the acquired data; wherein, the storage area contains the category identifier of the new action;
[0062] Step 6: Extract video data and corresponding category identifiers from the storage area to train the action classification model 8, thereby enabling action classification to recognize new actions.
[0063] In this example scenario, the action classification model 8 can already recognize three types of actions: A, B, and C. However, the accuracy of these three types of actions is insufficient and needs improvement. For example, to improve the accuracy of recognizing type A actions, a folder named "Action A" can be created in advance on the hard drive built into server 11. This "Action A" folder is used to collect training data for type A actions. A test subject 10 continuously performs various forms of type A actions in front of the video capture device 12. It is easy to understand that even a simple action like drinking water can have multiple ways of representing the semantic meaning of drinking water, thus constituting various forms of type A actions. These type A actions are continuously captured by the camera. Correspondingly, the video data of the captured type A actions is input into server 11 for real-time recognition and classification. The classification results include correctly classified image frames and incorrectly classified image frames. The incorrectly classified image frames are stored in the "Action A" folder for training the action classification model 8, thereby improving the accuracy of A action recognition.
[0064] This invention provides a training data acquisition device, comprising:
[0065] The acquisition module is used to acquire video data of the test actions.
[0066] In this embodiment, a video capture device 12 can be set up, and the test subject 10 continuously performs existing actions within the shooting area of the video capture device 12. After the video capture device 12 acquires video data, it transmits the video data to the server 11 via the network. It should be noted that there can be multiple types of existing actions, but in one capture, the test subject 10 can only perform one action. In this embodiment, the existing action is an action that the action classification model can recognize.
[0067] The action detection and recognition module is used to input the video data of the test action into a trained object detection model to perform object detection and obtain object detection results; and to input the object detection results into a trained action classification model to perform action classification and obtain action classification results.
[0068] In this embodiment, the object detection model can be a YOLO system series object detection model, such as the YOLOv5 model. The YOLOv5 model can be obtained using training data, such as images of people, specifically images of people indoors, such as an office. In this embodiment, the preferred training data is images of people sitting in an office, facing the video capture device 12. The corresponding annotations are for people, that is, drawing frames around the people in the images.
[0069] In this embodiment, the action classification model can employ a convolutional neural network or a recurrent neural network, or a spatio-temporal attention network (STN) model. The training data for the action classification model can be obtained using data such as images of people, labeled with the action category of the person in the image. In this embodiment, the preferred training data for the action classification model is images of people in an office, with the person facing the video capture device 12. The corresponding label is the action category of the person in the image, which corresponds to a specific scenario. Specifically, the action category is a sensitive or prohibited action in a particular scenario. For example, in a scenario preventing data leakage, the action category could be the action of copying text, or the action of taking a picture of a computer screen, etc. These actions could lead to data leakage and are therefore sensitive actions. In a scenario preventing cheating on exams, the action category could also be the action of using a mobile phone. In some scenarios, the action category could also be the action of eating. For example, in service positions of large organizations, it is necessary to monitor the actions of employees in real time to prevent employees from performing prohibited actions that could affect the organization's image.
[0070] The storage and annotation module is used to store the action image in the preset storage area and automatically annotate it when the action classification result identifies an action image that does not belong to the preset storage area. The storage area is configured with category identifiers for existing actions, and the category identifiers correspond to the action category of the test action being performed.
[0071] In this embodiment, an identification error can be represented as the action category that should be identified as the existing action in step S1, but the action classification model fails to identify it as such. For example, the action classification model can identify three action categories: A, B, and C. Correspondingly, in step S1, the test subject 10 performs action A, but the model outputs B or C, or it cannot give a specific category. That is, the action classification model believes that the actions performed by the test subject 10 in step S1 do not belong to the three action categories A, B, and C.
[0072] In this embodiment, the storage area corresponding to the existing action is pre-configured. The storage area can be a folder, a storage area directly set in the local storage space of the server, such as the hard drive of the local server, or a storage area set in a remote repository, such as a file system (FS), a distributed file system (HDFS), a network file system (NFS), or network attached storage (NAS). It can also be stored in cloud storage, such as Amazon S3, Google Cloud Storage, etc.
[0073] This invention provides a training data acquisition system, comprising:
[0074] Video capture device 12 is used to continuously capture the actions being performed and upload the captured data;
[0075] In this embodiment, a video capture device 12 can be set up. The test subject 10 continuously performs existing actions within the shooting area of the video capture device 12, such as playing with a mobile phone, copying, eating, or other actions. After the video capture device 12 captures the video data, it transmits the video data to the server 11 via the network. It should be noted that there can be multiple types of existing actions, but the test subject 10 can only perform one action in a single capture. In this embodiment, the existing action is one that the action classification model can recognize.
[0076] Server 11 is used to receive collected data and identify action categories. Server 11 is equipped with a pre-trained object detection model and action classification model.
[0077] In this embodiment, after obtaining the video data, the server 11 sequentially inputs the video data into the pre-trained object detection model and action classification model to perform object detection and action classification.
[0078] In this embodiment, the object detection model can be a YOLO system series object detection model, such as the YOLOv5 model. The YOLOv5 model can be trained using data such as images of people, specifically images of people indoors, such as an office. In this embodiment, the preferred training data is images of people sitting in an office, facing the video capture device. The corresponding annotations are for people, that is, drawing bounding boxes around the people in the image.
[0079] In this embodiment, the action classification model can employ a convolutional neural network or a recurrent neural network, or a spatio-temporal attention network (STN) model. The training data for the action classification model can be obtained using data such as images of people, labeled with the action category of the person in the image. In this embodiment, the preferred training data for the action classification model is images of people in an office, with the person facing the video capture device 12. The corresponding label is the action category of the person in the image, which corresponds to a specific scenario. Specifically, the action category is a sensitive or prohibited action in a particular scenario. For example, in a scenario preventing data leakage, the action category could be the action of copying text, or the action of taking a picture of a computer screen, etc. These actions could lead to data leakage and are therefore sensitive actions. In a scenario preventing cheating on exams, the action category could also be the action of using a mobile phone. In some scenarios, the action category could also be the action of eating. For example, in service positions of large organizations, it is necessary to monitor the actions of employees in real time to prevent employees from performing prohibited actions that could affect the organization's image.
[0080] Client 9 provides a real-time recognition interface to display the actions performed and the recognition results given by the model in real time; and provides a confidence score area to display the confidence value in real time.
[0081] This could involve providing a web page with a real-time recognition interface, displaying the user's actions and the model's recognition results, i.e., the specific action categories. The web page could also include a confidence score area, displaying the specific confidence value in real time for the tester or other users to view.
[0082] The aforementioned server 11 can be an electronic device with certain computing capabilities. For example, server 11 can be a server in a distributed system, or a system with multiple processors, memory, network communication modules, etc., operating collaboratively. Server 11 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. Server 11 can also be a server cluster formed by several servers. Alternatively, with the development of science and technology, server 11 can also be a new technical means capable of realizing the corresponding functions of the embodiments described in the specification. For example, it can be a new form of "server" based on quantum computing.
[0083] The aforementioned client can be an electronic device with network access capabilities. Specifically, for example, the terminal can be a desktop computer, tablet computer, laptop computer, smartphone, etc. Alternatively, the terminal can also be software that can run on the electronic device.
[0084] The aforementioned network can be any type of network, which can use any of the various available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. One or more networks can be a Local Area Network (LAN), an Ethernet-based network, a Token Ring network, a Wide Area Network (WAN), the Internet, a Virtual Network, a Virtual Private Network (VPN), an Intranet, an Extranet, a Public Switched Telephone Network (PSTN), an Infrared Network, a Wireless Network (e.g., Bluetooth, Wi-Fi), and / or any combination of these and / or other networks.
[0085] The specific implementation of the system or device can be achieved by referring to the aforementioned method implementation method.
[0086] The present invention provides a server 11, including a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the steps of the training data acquisition method.
[0087] The present invention provides a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, it implements the steps of the training data acquisition method.
[0088] In summary, this invention trains only the action classification model 8, adding only action category-related training data to reduce the overall training cost. Secondly, to achieve rapid training data acquisition and automatic annotation, this technical solution sets up corresponding training data storage areas for existing action categories or action categories to be added. When acquiring image data of actions, different image data are stored in the aforementioned training data storage areas, achieving rapid training data acquisition and automatic annotation. During training, training data is extracted from the aforementioned training data storage areas to train the model, enabling the addition of action categories to be detected and further improving the detection accuracy of existing action categories. This addresses rapidly changing action recognition needs at extremely low cost.
[0089] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0090] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0091] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0092] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0093] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art, under the guidance of the present invention, can make many modifications without departing from the spirit and scope of the claims, and all such modifications are within the protection scope of the present invention.
[0094] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for collecting training data, characterized by, The method comprises the following steps: acquiring video data of a test action; inputting the video data of the test action into a trained object detection model to perform object detection and obtain an object detection result; inputting the object detection result into a trained action classification model to perform action classification and obtain an action classification result; when the action classification result identifies an action image that does not belong to a preset storage area, storing the action image in the preset storage area and performing automatic labeling, including: when an error video frame is identified, the video frames before and after the time sequence reflected by the video data have similar data characteristics with the error video frame, taking the identified error image frame as a starting point, searching for a plurality of image frames according to a preset rule within a certain time window, and putting them into the storage area; the storage area is configured with a category identifier of an existing action, the category identifier corresponds to an action category of the executed test action, and the storage area stores a decomposed action representation of the action category.
2. The training data collection method of claim 1, wherein, When the action classification model needs to add the recognition of a new action, the acquisition method further comprises: acquiring video data of the same new action, the new action being an action of an action category that does not exist in the training data of the action classification model; storing the video data of the new action collected in the storage area corresponding to the new action and performing automatic labeling, the storage area being provided with a category identifier of the new action; extracting the video data and the corresponding category identifier from the storage area to train the action classification model, so as to recognize the new action.
3. The method of claim 1, wherein, The same existing action is continuously executed during the acquisition process.
4. The method of claim 1, wherein, The object detection model uses a YOLO series object detection model.
5. The method of claim 1, wherein, The action classification model uses a convolutional neural network, a recurrent neural network or a Spatio-Temporal Attention Network model.
6. The method of claim 1, wherein, The acquisition method further comprises: displaying and evaluating the confidence of the actual executed action and the recognized action in real time.
7. An apparatus for collecting training data, characterized by The method comprises the following steps: an acquisition module for acquiring video data of a test action; an action detection and recognition module for inputting the video data of the test action into a trained object detection model to perform object detection and obtain an object detection result; inputting the object detection result into a trained action classification model to perform action classification and obtain an action classification result; a storage and labeling module for storing the action image in a preset storage area and performing automatic labeling when the action classification result identifies an action image that does not belong to a preset storage area, including: when an error video frame is identified, the video frames before and after the time sequence reflected by the video data have similar data characteristics with the error video frame, taking the identified error image frame as a starting point, searching for a plurality of image frames according to a preset rule within a certain time window, and putting them into the storage area; the storage area is configured with a category identifier of an existing action, the category identifier corresponds to an action category of the executed test action, and the storage area stores a decomposed action representation of the action category.
8. A system for collecting training data for implementing the method of collecting training data according to any one of claims 1 to 6, characterized in that, The method comprises the following steps: a video acquisition device for continuously acquiring an executed action and uploading the acquired data; a server for receiving the collected data and performing action category recognition, the server being provided with a pre-trained object detection model and an action classification model; a client for providing a real-time recognition interface to display the performed action and the recognition result given by the model in real time; and a confidence score area for displaying the value of the confidence score in real time.
9. A server, characterized by A computer program product comprising a processor, a memory, and a computer program stored on the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the steps of the method for collecting training data according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that, A computer readable storage medium having stored thereon a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for collecting training data according to any one of claims 1 to 6.
Citation Information
Patent Citations
Automatic driving control method, device and equipment, motor vehicle and storage medium
CN112465685A
Foreign matter detection method and system for tobacco rolling and packaging equipment
CN112766141A
Classification model determination method and related device
CN114897076A