Method for generating a video action recognition model and electronic device

CN115527080BActive Publication Date: 2026-08-07ALIBABA EAST CHINA CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA EAST CHINA CO LTD
Filing Date
2022-09-09
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]其中,视频动作识别模型需要通过训练的方式生成,在对这种视频动作识别模型进行训练时,需要获取一些视频数据作为训练样本,现有技术中,还需要通过人工的方式分别对各个视频数据进行观看,并按照动作对这些视频数据切分为动作片段,还可能需要进行标注等,因此,人工操作成本很高

Benefits of technology

[0045] The solution provided in this application allows for unsupervised training of a video action recognition model. If unsupervised training is involved, the original, unsegmented video data can be used directly. This means that specific input data is constructed by directly sampling video frames from the original video data. Since segmentation or cropping of the original video data is unnecessary, the manual operation cost during model training can be reduced. Furthermore, during each sampling, the sampling range in the time dimension can be determined based on the average number of frames per action. This significantly reduces the probability of input data crossing action categories, thus ensuring the effectiveness of the model's input data to a certain extent.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115527080B_ABST
    Figure CN115527080B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method for generating a video action recognition model and an electronic device. The method comprises: obtaining training sample data, the training sample data comprising multiple pieces of video data; in a process of training the video action recognition model, for an unsupervised training part, generating input data of the model based on a manner of sampling video frames from original video data which is not divided into action segments, wherein at each sampling time, a sampling range in a time dimension is determined according to an average duration frame number of a single action. Through the embodiments of the present application, the manual processing cost of the training sample data can be saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video action recognition technology, and in particular to methods and electronic devices for generating video action recognition models. Background Technology

[0002] Video understanding and recognition is one of the fundamental tasks of computer vision, and video action recognition is a challenging yet highly valuable task in practical applications. Video action recognition refers to identifying the specific types of actions present in a video using a video action recognition model.

[0003] Among them, video action recognition models need to be generated through training. When training such video action recognition models, some video data needs to be obtained as training samples. In the existing technology, each video data needs to be viewed manually and segmented into action segments according to the actions. Labeling may also be required. Therefore, the manual operation cost is very high. Summary of the Invention

[0004] This application provides a method and electronic device for generating video action recognition models, which can save the cost of manual processing of training sample data.

[0005] This application provides the following solution:

[0006] A method for generating a video action recognition model includes:

[0007] Acquire training sample data, which includes multiple video data sets;

[0008] In the process of training the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined based on the average number of frames of a single action.

[0009] This also includes:

[0010] When constructing the same set of differentiated input data by sampling the same video data at least twice, each video frame sampling has the same sampling start point and the same sampling range on the time axis, so as to evaluate the training effect of the algorithm model through multiple sets of differentiated input data constructed.

[0011] In the process of sampling the same video data at least twice, one video frame sampling uses uniform sampling to determine the frame interval, while the other video frame samplings use an offset added to the uniform sampling to determine the sampling frame interval.

[0012] The offset can be a fixed value or a random value.

[0013] This also includes:

[0014] Representative video data is selected from the multiple video datasets to be segmented into multiple action segments, and action category annotation information for each action segment is obtained so that the algorithm model can be trained using a combination of supervised and unsupervised methods.

[0015] The step of selecting a representative portion of video data from the plurality of video data includes:

[0016] Select portions of video data from the multiple video datasets, each representing a different time period and a different person.

[0017] The step of selecting a representative portion of video data from the plurality of video data includes:

[0018] The algorithm model is used to predict the motion segments and corresponding motion categories contained in the video data;

[0019] From multiple action segments corresponding to the same action category, identify the action segments that differ significantly from each other, and use them as representative action segments for the corresponding action category;

[0020] The video data containing the representative action segments are selected as representative video data to participate in supervised training.

[0021] The step of identifying the action segments with significant differences from multiple action segments corresponding to the same action category includes:

[0022] The algorithm model is used to obtain feature vectors of multiple action segments under the same action category;

[0023] The difference quantization value between different action segments is determined by calculating the distance between the feature vectors of the multiple action segments;

[0024] Based on the quantified difference values ​​between the different action segments, the action segments with significant differences from each other are identified.

[0025] A video processing method, comprising:

[0026] Video data is obtained by capturing videos of the production process of workers in the target factory.

[0027] The pre-generated video action recognition model is used to identify the types of actions performed by workers in the video data; wherein, when training the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames of the original video data that has not been segmented according to action segments, wherein, in each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action;

[0028] Based on the identified action type and the action type specification information required for the production process in which the worker is located, it is determined whether the actions performed by the worker during the production process comply with the specifications.

[0029] This also includes:

[0030] If the worker's actions during the production process do not conform to the specifications, a prompt message is sent to the worker.

[0031] This also includes:

[0032] If the actions performed by the worker during the production process do not conform to the specifications, the corresponding production object identification information is determined, and quality inspection prompt information is provided to the corresponding quality inspection user.

[0033] An apparatus for generating video action recognition models, comprising:

[0034] A training sample data acquisition unit is used to acquire training sample data, which includes multiple video data sets.

[0035] The input data construction unit is used to generate the model's input data during the training of the video action recognition model. For the unsupervised training part, it is based on sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action.

[0036] A video processing apparatus, comprising:

[0037] The video acquisition unit is used to capture video of the production process of workers in the target factory to obtain video data;

[0038] An action recognition unit is used to identify the type of action performed by a worker in the video data using a pre-generated video action recognition model. When training the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames of the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action.

[0039] The judgment unit is used to determine whether the actions performed by the worker in the production process conform to the specifications based on the identified action type and the action type specification information required for the production link where the worker is located.

[0040] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any of the preceding methods.

[0041] An electronic device, comprising:

[0042] One or more processors; and

[0043] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the preceding descriptions.

[0044] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0045] The solution provided in this application allows for unsupervised training of a video action recognition model. If unsupervised training is involved, the original, unsegmented video data can be used directly. This means that specific input data is constructed by directly sampling video frames from the original video data. Since segmentation or cropping of the original video data is unnecessary, the manual operation cost during model training can be reduced. Furthermore, during each sampling, the sampling range in the time dimension can be determined based on the average number of frames per action. This significantly reduces the probability of input data crossing action categories, thus ensuring the effectiveness of the model's input data to a certain extent.

[0046] Furthermore, when constructing the same set of differentiated input data by sampling the same video data at least twice, each video frame sampling can have the same sampling start point and the same sampling range on the time axis. This allows the training effect of the algorithm model to be evaluated using multiple sets of differentiated input data. In this way, the requirement that differentiated input data within the same set correspond to the same or similar action content can be met, thus enabling the constructed multiple sets of differentiated input data to be used to evaluate the model training effect.

[0047] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0048] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0049] Figure 1 This is a schematic diagram of the training sample processing method in the existing technology;

[0050] Figure 2 This is a schematic diagram of the training sample processing method provided in the embodiments of this application;

[0051] Figure 3 This is a schematic diagram of the sampling method provided in the embodiments of this application;

[0052] Figure 4 This is a flowchart of the first method provided in the embodiments of this application;

[0053] Figure 5 This is a flowchart of the second method provided in the embodiments of this application;

[0054] Figure 6 This is a schematic diagram of the first device provided in the embodiments of this application;

[0055] Figure 7 This is a schematic diagram of the second device provided in the embodiments of this application;

[0056] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0057] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0058] To facilitate understanding of the solutions provided in the embodiments of this application, the algorithm model training methods in the prior art will be introduced first.

[0059] For algorithm model training, there are generally three methods: supervised training, unsupervised training, and semi-supervised training. Video action recognition models are no exception. They will be briefly introduced below.

[0060] Supervised training refers to a method where each input data point has corresponding annotation information, which serves as the theoretical result for model training. In other words, supervised training requires all input data to undergo a data annotation process. If a video action recognition model is trained entirely using supervised training, a large amount of action video data needs to be collected and fully annotated. Specifically, annotation requires annotators to manually view complete video data and annotate the start and end times of multiple action segments, as well as their corresponding action categories. This allows a video data point to be segmented into multiple action segments, each of which may contain only one complete action and its corresponding action category annotation information. While this method can yield relatively accurate prediction results, the high cost due to the significant manual labor involved in the annotation process makes it costly. Furthermore, models trained in this way tend to have relatively poor generalization ability. Generalization ability refers to the difference between the model's accuracy when processing unseen data and its accuracy when processing training data. The smaller this difference, the stronger the model's generalization ability.

[0061] II. Unsupervised training refers to training a model based on the differences between the relationships between several input data points and the relationships between their corresponding model outputs, without known correct annotation information for the input data. If unsupervised training is used to train a video action recognition model, it means that specific action categories do not need to be labeled, reducing manual costs. Furthermore, compared to supervised training, the model can achieve stronger generalization. However, since an original video is usually composed of multiple action segments connected together on a timeline, and the algorithm model is trained under the premise that each input data point contains only one action, in existing technologies, when training a video action recognition model using unsupervised methods, it is still necessary to segment the original video data into individual action segments, with each segment containing only one complete action. Moreover, during the segmentation of the original video data, it is still necessary to manually mark the start and end times of each segment. In other words, compared to the labeled data in supervised methods, only the step of labeling the specific action categories corresponding to the action segments is omitted. However, in the overall process of action video annotation, action segmentation is a more time-consuming step that requires manual judgment. When performing action segmentation, the annotators have already watched the entire action segment and know the action type. Not performing action labeling only saves a simple manual action of point selection, while the time required to watch the entire video for action segmentation cannot be avoided.

[0062] Third, semi-supervised training refers to a combination of supervised and unsupervised training. That is, a portion of the training data is labeled, while the majority of the training data is unlabeled. Specifically, labeled and unlabeled data can be mixed in a certain proportion; for example, labeled data can be no more than 10%, such as 10% labeled data and 90% unlabeled data, etc. During training, the two methods can be alternated or performed simultaneously. For example, among the multiple data points input into the model at one time, some can be labeled data and the other part can be unlabeled data, etc. This semi-supervised approach can combine the advantages of supervised and unsupervised training, achieving higher prediction accuracy and stronger generalization capabilities at a lower cost.

[0063] However, in the process of training the video action recognition model using the semi-supervised method described above, since unsupervised training is still involved, if the existing technology is used, the specific video data still needs to be segmented to obtain multiple action segments to ensure that each action segment contains only one complete action. That is, as... Figure 1 The diagram illustrates the data processing flow for semi-supervised training as described above. Specifically, the supervised training portion requires segmenting the video data into multiple action fragments and labeling each fragment with specific action category information. The unsupervised training portion similarly involves segmenting the video data into multiple action fragments, but without labeling specific action categories. Therefore, this semi-supervised training method still suffers from the problem of high labor costs associated with the segmentation process.

[0064] In view of the above situation, in the embodiments of this application, a semi-supervised training method can be adopted. At the same time, the data input format of the unsupervised training part is improved to reduce the training cost. Of course, this improvement method is also applicable to the method of completely adopting unsupervised training.

[0065] The specific improvement involves using unsegmented, raw video data as input for unsupervised training. Since segmentation is unnecessary, there's no need for manual review of each video beforehand, nor is it necessary to add start and end time markers, significantly reducing labor costs. For example... Figure 2 As shown in the embodiments of this application, if a semi-supervised training method is adopted, a small portion of the video data can be segmented and labeled, while the majority of the data does not need to be segmented or labeled and can be directly used in the training of the model.

[0066] Of course, since the raw, unsegmented video data may be composed of multiple action segments connected together, it does not conform to the assumption during model training that each input data point contains only one action. To address this, considering that the construction of input data also involves video frame sampling (which requires each input data point to have a consistent length, i.e., the same number of frames, and thus can be achieved through sampling), the specific sampling method can be improved to better satisfy the aforementioned assumption.

[0067] Specifically, regarding the video frame sampling process, we will first take the existing technology of pre-segmenting action segments as an example. Action recognition refers to processing a video segment (i.e., an action fragment) as input to a recognition model, which then outputs a classification result of the action category within that segment. Specifically, while the algorithm model can process multiple input data at once, classification tasks typically require a fixed dimension (i.e., length; in this embodiment, this could be the number of video frames contained in each input data point) for a single input data point. However, for action fragments, the number of video frames contained in different action fragments is usually inconsistent. Therefore, it is necessary to first sample video frames for each action fragment, that is, to sample a fixed number of video frames from the complete action fragment using a preset sampling method to represent that action fragment and input them into the model for processing.

[0068] One prerequisite for training the model is that a single input data set typically contains only one type of action. In existing technologies where action segments are pre-segmented, each input data set is obtained by sampling video frames from that action segment, thus satisfying this prerequisite. However, in this embodiment, the sampling is performed on the unsegmented raw video, which contains various types of action content. Therefore, without control, a single input data set could contain multiple different action contents. To address this, this embodiment improves the sampling method. Specifically, the sampling range can be limited based on the average duration of a single action. For example, assuming the average duration of a single action is 25 frames, each sampling can be performed within a 25-frame range. For instance, if the input data is 8 frames long and a sampling start point is the 10th frame of the original video, then this sampling can extract 8 video frames between the 10th and 35th frames of the video as a single input data set, and so on. Since a single action typically lasts around 25 frames (a hypothetical value used for illustrative purposes only), sampling within this range significantly reduces the probability of input data spanning multiple action categories. In other words, this method ensures that the constructed input data contains only one action category, and even when actions cross categories, it minimizes the number of categories to two, thus guaranteeing the validity of most of the input data.

[0069] Furthermore, in unsupervised training, since the true labels of the input data are unknown, it is usually necessary to perform at least two different transformations on the same action segment to construct a set of differentiated input data in order to evaluate the training effect of the model. Then, the model is trained by maximizing the similarity between the corresponding two outputs based on the constraint that the differentiated data within the same set actually belong to the same category. Specifically, in self-supervised training for action recognition, it is necessary not only to transform the image content of each sampled video frame, but also to transform the temporal dimension covered by the sampled video frames through methods such as different sampling methods and sampling parameters. For example, ... Figure 3As shown in (A), assuming that 31 represents an action segment, two different input data sets (i.e., a set of differentiated input data) can be constructed by sampling this action segment twice. These two different input data sets have the same number of video frames, but their specific video frame compositions are different. For example, the first input data set might be frames 1, 4, 7, and 10 of a certain action segment, while the second might be frames 3, 6, 9, and 12, and so on. However, since they are sampled from the same action segment, if the model's training effect is good, the output results of the model after inputting the above two input data sets should be the same, or at least have a high degree of similarity. Therefore, the similarity between the model output results corresponding to the two differentiated input data sets constructed from the same action segment can be used to construct a loss function. During model training, the function value of this loss function is maximized by continuously adjusting the parameter values ​​in the model until the algorithm converges.

[0070] In the case of pre-segmenting multiple action segments, since the input data is constructed based on the specific action segments, a fixed interval combined with a random initial position can be used when sampling the same action segment each time. For example, in the aforementioned... Figure 3 In the example shown in (A), the two different sampling operations can both be spaced 3 frames apart, where the first sampling starts from frame 1, the second sampling starts from frame 3, and so on.

[0071] Specifically, in this embodiment, since training is performed directly using unsegmented raw video data, and the lengths of different raw video data vary, sampling of this raw video data is also necessary. However, because the raw video data has not been segmented, it may include multiple different action segments, as well as some non-action segments, etc. For example, as... Figure 3 As shown at point 32 in (B), action segments and non-action segments are represented by different grayscale blocks. Currently, this illustration only shows the difference between action segments and non-action segments, and does not show the difference between different action segments. However, it can be understood that each different action segment in the same video data usually corresponds to a different action.

[0072] In the above Figure 3 In the case shown in (B), if still using Figure 3As shown in (A), the sampling method of fixed intervals combined with random initial positions may result in excessively large differences between a set of differentiated data constructed from two (or more) samplings. That is, the actual action content contained in a fixed number of video frames obtained from two samplings may differ significantly. This renders the prerequisite of constructing differentiated change results on the "same data" required for unsupervised training invalid, leading to poorer recognition results. In other words, as... Figure 3 As shown in (B), the two diagrams illustrate two sampling operations based on the same video data. The arrows indicate the positions of the sampled video frames within the video. It can be seen from the diagrams that because the starting positions of the two sampling operations are random, the first sampling primarily extracts video frames from the first action segment, the second non-action segment, and some frames from the second action segment. The second sampling primarily extracts video frames from the second and third action segments, and so on. This results in two input data pairs corresponding to data frames from different action segments. Clearly, such input data pairs are unsuitable for evaluating the model's training performance.

[0073] To address the above situation, in this embodiment of the application, in addition to unsupervised training using unsegmented raw video data, a differentiated data construction method of "fixed sampling starting point + fixed sampling range" can be adopted to construct multiple sets of differentiated input data that can be used to effectively evaluate the model training effect. That is, specifically, in the process of constructing a set of differentiated input data by sampling at least twice based on the same video data, each sampling can have the same sampling starting point and the same sampling range on the time axis. For example, as... Figure 3 As shown in (C), assuming sampling starts from frame 10 of the video and samples from video frames within a range of 25 frames after that frame, this sampling method fully considers the characteristics of the distribution of unknown effective action segments in unsegmented videos. By fixing the sampling starting point and the sampling range, it limits the coverage of each sampling frame, thereby maximizing the similarity of action content contained in the sampled frames. Of course, to ensure the difference between each sampling result, it can be achieved by limiting the frame interval of each sampling. For example, one video frame sampling can use uniform sampling to determine the frame interval, while other video frame sampling can use an offset added to the uniform sampling to determine the sampling frame interval. Here, the offset can be random or a fixed value, etc. In addition, the aforementioned sampling range can also be determined based on the average number of frames of a single action. For example, assuming that the number of frames of a single action is usually 25, then the sampling range can be set to 25 frames, etc., which can reduce the probability of the same input data spanning multiple action categories.

[0074] In summary, through the embodiments of this application, when training a video action recognition model, if unsupervised training is involved, training can be directly performed based on the unsegmented raw video data. However, to ensure the effectiveness of the input data, the sampling range for each sample can be limited. Specifically, this sampling range can be determined based on the average duration of a single action, thereby reducing the probability of cross-action categories in the obtained input data. Furthermore, in the process of constructing multiple sets of differentiated input data for evaluating the training effect of the algorithm model by sampling video data, when multiple samples are performed based on the same video data to construct the same set of differentiated input data, each video frame sample can have the same sampling starting point and the same sampling range on the time axis. In this way, the similarity of the action content contained in the input data obtained from multiple samplings is maximized, thereby enabling effective evaluation of the model training results.

[0075] The specific implementation schemes provided in the embodiments of this application will be described in detail below.

[0076] Example 1

[0077] First, Embodiment 1 of this application provides a method for generating a video action recognition model, see [link to embodiment]. Figure 4 The method may include:

[0078] S401: Obtain training sample data, which includes multiple video data sets.

[0079] Specifically, since the embodiments of this application mainly focus on training a video action recognition model, the specific training samples can primarily be video data including specific action content. Furthermore, the specific video action recognition model can be specifically trained for a particular scenario. For example, as described in the background section, if it is necessary to recognize the actions of workers in a smart garment factory, the specific training samples could be video data collected from the workers' work processes in such a smart garment factory. Specifically, new video data may be generated daily by each individual, and this video data can all be used as training samples in the embodiments of this application to participate in the training of the video action recognition model.

[0080] In this embodiment, a semi-supervised training method is primarily employed. This semi-supervised training process can include both supervised and unsupervised training components. Therefore, after obtaining specific training samples, a portion can be allocated for supervised training. This portion needs to be pre-segmented into multiple action segments, ensuring each segment contains only one action. Furthermore, the action segments can be categorized, i.e., labeled with specific action categories. The remaining training samples for unsupervised training do not require segmentation or labeling.

[0081] It should be noted that, in semi-supervised training, although a combination of supervised and unsupervised methods is involved, the specific algorithm model can be the same. That is, the same algorithm model can be trained using both supervised and unsupervised methods. During the training process, supervised and unsupervised methods can be performed alternately or simultaneously, and so on. Of course, the solution provided in this application can also be applied to purely unsupervised training processes.

[0082] S402: During the training of the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined based on the average number of frames of a single action.

[0083] In this embodiment, the unsupervised training can be performed directly based on the unsegmented raw video data. Therefore, when constructing specific input data for the algorithm model, video frame sampling is performed directly based on this raw video data to ensure that each input data has a fixed dimension. To ensure the validity of the input data, the sampling range can be limited. Specifically, in this embodiment, the sampling range can be set to a relatively small range. This range can be determined based on the average duration of a single action to reduce the probability that the same input data contains multiple categories of different actions. That is, as mentioned earlier, if an input data set includes multiple different categories of actions, the probability that the model identifies the input data as belonging to each action category may be relatively low. Therefore, the number of actions spanned by the input data can be minimized. This can be achieved by reducing the sampling range. For example, assuming the average duration of a single action is 25 frames, the sampling range can be set to 25. That is, assuming the Nth frame is the starting point, sampling can be performed within the range from the Nth frame to the N+25th frame, thereby reducing the probability that the same input data spans multiple different action categories.

[0084] It should be noted that since an original video may be quite long, and this application does not restrict the sampling start point position, meaning the exact location of the sampling start point is unknown, controlling the sampling range within the average duration of a single action can reduce the probability of the sampling results spanning multiple action categories. For example, if the sampling start point happens to be at the beginning of an action segment, this is ideal; by controlling the sampling range, the sampled input data can be made to include only a single action content as much as possible. However, if the sampling start point is located in the middle of an action segment, controlling the sampling range within the duration of a single action may result in two different action contents within the same input data. Alternatively, if the sampling start point is located near the end of an action segment, and the next action segment is relatively short, three or more different action contents may appear within the same input data, and so on. Of course, in practical applications, limiting the sampling range significantly reduces the probability of the third scenario mentioned above. In other words, since the average number of frames for a single action can usually be calculated, and there may be some non-action segments interspersed between different action segments in the same video, overall, if any point is taken as the sampling starting point, the probability of including multiple different action segments within a certain sampling range after that point will be relatively low. In most cases, it can be controlled within the range of one or two action contents.

[0085] In cases where the same input data spans two categories, if the model is trained well, although the probability of identifying the input data as belonging to a certain action category may only be 80% (because 20% of the video content belongs to another action category), this identification result still has some value compared to the higher probability identification result for a single action, and can play a positive role in the model's learning and training. Of course, if the input data constructed using the method provided in this application results in a low probability of the input data belonging to each action category, it may be due to the input data containing action content from multiple categories. In this case, such data (usually only a small portion) can be removed to avoid negatively impacting the accuracy of the model training results.

[0086] In addition, in order to construct a loss function that can evaluate the training effect of unsupervised training and update the parameters in the model accordingly, as mentioned above, it is necessary to construct multiple sets of differentiated input data. Specifically, the same video data can be sampled at least twice to construct a set of differentiated input data based on the same or similar action content. Then, the training effect of the model can be evaluated based on the similarity between the action category results predicted by the model for the same set of differentiated input data.

[0087] In the process of constructing differentiated input data described above, in this embodiment of the application, since the specific original video data may include various types of action content, a sampling method with a fixed sampling starting point and a fixed sampling range can be adopted. That is, when constructing a set of differentiated input data based on the same video data, each sampling can start from the same starting point and sample within a certain range. Of course, in order to ensure that there are differences between each sampling result, different sampling intervals can be set. For example, in a specific implementation, one video frame sampling can use uniform sampling to determine the frame interval, while other video frame sampling can use a method of adding an offset to the uniform sampling to determine the sampling frame interval. The specific offset can be a fixed value, or it can be a random value, etc.

[0088] For example, assuming the sampling starting point is the first frame, the range is 25 frames, and a total of 8 frames need to be sampled, then under uniform sampling, the frame interval can be 3 frames. That is, the sampling results can be frames 1, 4, 7, 10, 13, 16, 19, and 22. The second time, an offset can be added to the above, for example, shifting backward by one frame, so the sampling results can be frames 1, 5, 8, 11, 14, 17, 20, and 23, and so on. Alternatively, the specific offset can also be random; for example, some might shift forward, some backward, and there could even be some cases where the offset is 0, and so on.

[0089] It's also worth noting that in practice, since a single piece of raw video data can often be quite long (relative to a single action segment), multiple sets of differentiated input data can be constructed from it. Each set of differentiated input data can include at least two data points, as long as the same set of differentiated input data is obtained by sampling from the same sampling starting point and within the same sampling range. For example, assuming the first frame of a video data is used as the sampling starting point, and sampling is performed twice within the following 25 frames (i.e., frames 1 to 25), one set of differentiated input data can be obtained. Similarly, the 30th frame of the same video data can be used as the sampling starting point, and sampling is performed twice within the following 25 frames (i.e., frames 30 to 55), resulting in another set of differentiated input data, and so on.

[0090] The above describes the unsupervised training method. In practice, it is often combined with supervised training for semi-supervised training. Therefore, as mentioned earlier, this involves segmenting and labeling a portion of the training samples. Specifically, a portion of the video data can be selected from the multiple video datasets to segment into multiple action segments, and action category labeling information for each action segment can be obtained. This allows the algorithm model to be trained using a combination of supervised and unsupervised methods.

[0091] Specifically, in selecting video data for supervised training, although only a small portion of the video data can be chosen, to improve the model's training effect and efficiency, the quality of the training samples should be maximized. That is, representative video data should be selected for supervised training. For instance, assuming the algorithm needs to identify five specific action categories, the training samples (action segments) can include action segments corresponding to each of these five categories, with multiple action segments under each category. Even within the same action category, different people may perform the action differently, or different scenarios may lead to different actions, and so on. Therefore, to improve model training performance, multiple action segments within the same action category should cover these different approaches as much as possible, thereby improving the model's generalization ability. In other words, action segments that demonstrate multiple different approaches within the same action category are more meaningful for model training. Therefore, the selection of training data (i.e., choosing which training samples to segment and label) can be controlled in some ways to improve the quality of the samples participating in supervised training, thereby improving the model's training efficiency.

[0092] To achieve the above objectives, a relatively simple approach is to select portions of video data from multiple video datasets, each representing different times and individuals. In other words, the samples can be distributed across different times and individuals. Since the probability of differences in the specific actions performed by different individuals at different times on the same type of action is relatively high, this method can effectively achieve the desired outcome.

[0093] Alternatively, the algorithm model can be trained unsupervised to develop predictive capabilities. This model can then be used to predict action segments and their corresponding action categories within the training video data. Specifically, the algorithm can segment the video data for identification. For example, a certain number of video frames can be extracted from a 2-second video segment, and the action categories within each segment can be identified. Adjacent video segments of the same category can be connected to obtain specific action segments (though, in unsupervised training, the names of specific action categories cannot be directly given). Then, from multiple action segments corresponding to the same action category, the most significantly different action segments can be identified as representative action segments for that category. The video data containing these representative action segments is then used as part of the supervised training data.

[0094] Specifically, regarding the differences between multiple action segments corresponding to the same action category, since the algorithm model can generate corresponding feature vectors for specific action segments during the action recognition process, the algorithm model can be used to obtain the feature vectors of multiple action segments included in the same action category. Then, the difference quantification value between different action segments can be determined by calculating the distance between the feature vectors of multiple action segments. Furthermore, based on the difference quantification value between different action segments, the action segments with larger differences can be identified.

[0095] In other words, after training the model in an unsupervised manner, action segments can be identified from the raw video data. Although the specific category names of the action segments are unknown, it is possible to roughly determine which action segments belong to the same category and generate specific feature vectors for each action segment. Thus, the degree of difference between the action content contained in each action segment can be determined based on the distance between the feature vectors of different action segments within the same category. Since the specific action segments belong to the raw video data, this method can roughly determine which video data contains more representative action segments. These video data can be selected, and then manually segmented and labeled to improve the quality of the training data used in supervised training.

[0096] It's important to note that after training the video action recognition model, it can be applied to identify or predict action categories within videos. The target of this prediction can be a pre-recorded video file or a real-time video stream. Regardless of whether it's a video file or a real-time stream, the prediction process can involve segmenting the prediction into segments and then aggregating the action category predictions for each segment. For example, adjacent segments of the same category can be connected to form an action fragment. During segmented prediction, a 2-second interval (or other values) can be used as a segment, and video frames can be sampled from each 2-second segment and input into the model for prediction. The prediction or recognition results for action fragments can then be used to guide or correct worker operating procedures, thereby improving the quality of manufactured products.

[0097] In summary, through the embodiments of this application, when training a video action recognition model, if unsupervised training is involved, unsupervised training can be directly performed using the raw, unsegmented video data. That is, specific input data can be constructed by directly sampling video frames from the raw video data. This reduces the manual operation cost during model training because it eliminates the need for segmentation or cropping of the raw video data. Furthermore, during each sampling, the sampling range in the time dimension can be determined based on the average duration of a single action. This significantly reduces the probability of input data crossing action categories, thus ensuring the effectiveness of the model's input data to a certain extent.

[0098] Furthermore, when constructing the same set of differentiated input data by sampling the same video data at least twice, each video frame sampling can have the same sampling start point and the same sampling range on the time axis. This allows the training effect of the algorithm model to be evaluated using multiple sets of differentiated input data. In this way, the requirement that differentiated input data within the same set correspond to the same or similar action content can be met, thus enabling the constructed multiple sets of differentiated input data to be used to evaluate the model training effect.

[0099] Example 2

[0100] This second embodiment primarily introduces the application of specific video action recognition models. Specifically, video action recognition has a wide range of applications. For example, in smart factories or digital factories (such as smart garment factories), workers at each stage perform their respective actions according to preset standards to collectively complete production tasks. Therefore, the standardization of workers' actions at each stage often determines the quality of the final product. In this process, the production process can be managed digitally to improve product quality. This requires generating effective video action recognition models to automatically identify whether workers' actions conform to standards from recorded videos or real-time video streams, and so on.

[0101] Specifically, Embodiment 2 of this application provides a video processing method for application in the aforementioned smart factory or digital factory, see [link to embodiment]. Figure 5 The method may include:

[0102] S501: Collect video footage of the production process of workers in the target factory to obtain video data.

[0103] The video data can be recorded video files or real-time captured video streams. In practice, video capture devices can be deployed within the work area of ​​the target factory to capture video of workers' production processes. That is, the video data can include the specific actions performed by workers during the production process.

[0104] S502: Use a pre-generated video action recognition model to identify the type of action performed by the worker in the video data; wherein, when training the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames of the original video data that has not been segmented according to action segments, wherein, in each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action.

[0105] After identifying the video data to be processed, a pre-generated video action recognition model can be used to identify the types of actions performed by workers in the video data. This model can be established using the method described in Embodiment 1. Specifically, when identifying action types, as described in Embodiment 1, the video data can be segmented for prediction, and then the action category prediction results corresponding to each segment can be aggregated. For example, adjacent segments of the same category can be connected together to form an action clip, and so on. When performing segmented prediction, 2 seconds (or other values) can be used as a segment, and video frames can be sampled from each 2-second segment and input into the model for prediction, and so on.

[0106] S503: Based on the identified action type and the action type specification information required for the production process in which the worker is located, determine whether the action performed by the worker in the production process conforms to the specification.

[0107] To determine whether workers' actions are standardized, the standardized information of the action types required for each production stage can be pre-stored. After identifying the specific action type performed by a worker in the corresponding production stage, the algorithm can compare it with the standardized information to determine if the worker's actions conform to the standards. For example, in a production stage requiring the sequential execution of actions 1, 2, and 3, if a worker's execution of action 2 is not standardized, the algorithm may fail to recognize action 2. That is, the worker only performed actions 1 and 3 correctly, but not action 2, or performed it incorrectly. This could potentially affect the quality of the final product. Therefore, the judgment results in this embodiment can be used to guide or correct worker operation standards, thereby helping to improve the quality of the specific product.

[0108] In practice, factories can be equipped with large-screen displays, or individual workers can be equipped with terminal devices. When a worker's action is detected as non-compliant, a prompt message can be provided. For example, the prompt message can be displayed on a large screen or sent directly to the worker's individual terminal device, and so on.

[0109] Furthermore, if the actions performed by the worker during the production process do not conform to the specifications, the corresponding production object identification information can be determined, and quality inspection prompts can be provided to the relevant quality inspection user. For example, based on the time period in which the non-conforming operation was detected, the corresponding batch identifier of the production object can be determined. This batch identifier can then be provided to the quality inspection user to remind them to pay closer attention during the quality inspection of that batch of production objects, and so on.

[0110] For the parts of this embodiment that are not described in detail, please refer to the description in embodiment one, which will not be repeated here.

[0111] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0112] Corresponding to the aforementioned Embodiment 1, this application also provides an apparatus for generating a video action recognition model, see [link to embodiment 1]. Figure 6 The device may include:

[0113] The training sample data acquisition unit 601 is used to acquire training sample data, which includes multiple video data sets.

[0114] The input data construction unit 602 is used to generate input data for the unsupervised training part of the video action recognition model by sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined based on the average number of frames of a single action.

[0115] Specifically, the device may also include:

[0116] The differential input data construction unit is used to construct the same set of differential input data by sampling the same video data at least twice, wherein each video frame sampling has the same sampling start point and the same sampling range on the time axis, so as to evaluate the training effect of the algorithm model by constructing multiple sets of differential input data.

[0117] Specifically, in the process of sampling the same video data at least twice, one video frame sampling uses uniform sampling to determine the frame interval, while the other video frame sampling uses an offset added to the uniform sampling to determine the sampling frame interval.

[0118] The offset can be a fixed value or a random value.

[0119] In a specific implementation, the device may further include:

[0120] The sample selection unit is used to select a portion of representative video data from the multiple video data sets for segmenting into multiple action segments and obtaining action category annotation information for each action segment, so as to train the algorithm model through a combination of supervised and unsupervised methods.

[0121] Specifically, the sample selection unit can be used for

[0122] Select portions of video data from the multiple video datasets, each representing a different time period and a different person.

[0123] Alternatively, the sample selection unit can be specifically used for:

[0124] The prediction subunit is used to predict the action segments and corresponding action categories contained in the video data using the algorithm model.

[0125] The difference calculation subunit is used to identify the action segments with significant differences from multiple action segments corresponding to the same action category, and use them as representative action segments of the corresponding action category.

[0126] The sub-unit is selected to use the video data containing the representative action segment as the selected representative video data to participate in supervised training.

[0127] Specifically, the difference calculation subunit can be used for:

[0128] The algorithm model is used to obtain feature vectors of multiple action segments under the same action category;

[0129] The difference quantization value between different action segments is determined by calculating the distance between the feature vectors of the multiple action segments;

[0130] Based on the quantified difference values ​​between the different action segments, the action segments with significant differences from each other are identified.

[0131] Corresponding to the aforementioned Embodiment 2, this application also provides a video processing apparatus, see [link to embodiment 2]. Figure 7 The device may include:

[0132] The video acquisition unit 701 is used to acquire video of the production process of workers in the target factory to obtain video data;

[0133] The action recognition unit 702 is used to identify the type of action performed by the worker in the video data using a pre-generated video action recognition model; wherein, when training the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames of the original video data that has not been segmented according to action segments, wherein, in each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action.

[0134] The judgment unit 703 is used to determine whether the actions performed by the worker in the production process conform to the specifications based on the identified action type and the action type specification information required for the production link where the worker is located.

[0135] In a specific implementation, the device may further include:

[0136] The first prompting unit is used to send a prompting message to the worker if the worker's actions during the production process do not conform to the specifications.

[0137] Alternatively, it may also include:

[0138] The second prompting unit is used to determine the corresponding production object identification information and provide quality inspection prompt information to the corresponding quality inspection user if the actions performed by the worker during the production process do not conform to the specifications.

[0139] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0140] And an electronic device, comprising:

[0141] One or more processors; and

[0142] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0143] in, Figure 8An exemplary architecture of an electronic device is shown, which may include a processor 810, a video display adapter 811, a disk drive 812, an input / output interface 813, a network interface 814, and a memory 820. The processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and memory 820 can communicate with each other via a communication bus 830.

[0144] The processor 810 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to achieve the technical solution provided in this application.

[0145] The memory 820 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 820 can store the operating system 821 for controlling the operation of the electronic device 800, and the basic input / output system (BIOS) for controlling the low-level operations of the electronic device 800. Additionally, it can store a web browser 823, a data storage management system 824, and a model generation processing system 825, etc. The aforementioned model generation processing system 825 can be the application program that specifically implements the aforementioned steps in this embodiment. In summary, when implementing the technical solution provided in this application through software or firmware, the relevant program code is stored in the memory 820 and is called and executed by the processor 810.

[0146] The input / output interface 813 is used to connect input / output modules to enable information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0147] Network interface 814 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0148] Bus 830 includes a pathway for transmitting information between various components of the device, such as processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, and memory 820.

[0149] It should be noted that although the above-described device only shows the processor 810, video display adapter 811, disk drive 812, input / output interface 813, network interface 814, memory 820, bus 830, etc., in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the solution of this application, and does not necessarily include all the components shown in the figures.

[0150] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0151] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0152] The method and electronic device for generating video action recognition models provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for generating a video action recognition model, characterized in that, include: Acquire training sample data, which includes multiple video data sets; In the training process of the video action recognition model, for the unsupervised training part, the input data of the model is generated by sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined according to the average number of frames of a single action. When constructing the same set of differentiated input data by sampling the same video data at least twice, each video frame sampling has the same sampling starting point and the same sampling range on the time axis, and the number of sampling frames in each video frame sampling is the same. The frame interval is determined by uniform sampling in one video frame sampling, and the sampling frame interval is determined by adding an offset to the uniform sampling in other video frame sampling. In order to evaluate the training effect of the video action recognition model by constructing multiple sets of differentiated input data.

2. The method according to claim 1, characterized in that, The offset can be a fixed value or a random value.

3. The method according to claim 1, characterized in that, Also includes: Representative video data is selected from the multiple video datasets to be segmented into multiple action segments, and action category annotation information for each action segment is obtained, so as to train the video action recognition model in a combination of supervised and unsupervised methods.

4. The method according to claim 3, characterized in that, The step of selecting a representative portion of video data from the multiple video datasets includes: Select portions of video data from the multiple video datasets, each representing a different time period and a different person.

5. The method according to claim 3, characterized in that, The step of selecting a representative portion of video data from the multiple video datasets includes: The video action recognition model is used to predict the action segments and corresponding action categories contained in the video data; From multiple action segments corresponding to the same action category, identify the action segments that differ significantly from each other, and use them as representative action segments for the corresponding action category; The video data containing the representative action segments are selected as representative video data to participate in supervised training.

6. The method according to claim 5, characterized in that, The step of identifying the action segments with significant differences from multiple action segments corresponding to the same action category includes: The video action recognition model is used to obtain feature vectors of multiple action segments under the same action category; The difference quantization value between different action segments is determined by calculating the distance between the feature vectors of the multiple action segments; Based on the quantified difference values ​​between the different action segments, the action segments with significant differences from each other are identified.

7. A video processing method, characterized in that, include: Video data is obtained by capturing videos of the production process of workers in the target factory. A pre-generated video action recognition model is used to identify the types of actions performed by workers in the video data. During training of the video action recognition model, for the unsupervised training portion, input data for the model is generated by sampling video frames from the original video data that has not been segmented according to action segments. In each sampling, the sampling range in the time dimension is determined based on the average duration of a single action. When constructing the same set of differentiated input data by sampling the same video data at least twice, each video frame sampling has the same sampling starting point and the same sampling range on the time axis, and the number of sampling frames is the same each time. One video frame sampling uses uniform sampling to determine the frame interval, while other video frame sampling uses an offset added to the uniform sampling to determine the sampling frame interval. This allows for the evaluation of the training effect of the video action recognition model using multiple sets of differentiated input data. Based on the identified action type and the action type specification information required for the production process in which the worker is located, it is determined whether the actions performed by the worker during the production process comply with the specifications.

8. The method according to claim 7, characterized in that, Also includes: If the worker's actions during the production process do not conform to the specifications, a prompt message is sent to the worker.

9. The method according to claim 7, characterized in that, Also includes: If the actions performed by the worker during the production process do not conform to the specifications, the corresponding production object identification information is determined, and quality inspection prompt information is provided to the corresponding quality inspection user.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1 to 9.

11. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Video action recognition method based on time segmentation network and storage medium

    CN112733595A

  • Human body behavior recognition method and device

    CN114332693A