Sample data processing method and device applied to model training

By grouping and fusing sample videos to generate fused sample videos, and training the recognition model, the complexity problem of recognizing multi-category abnormal behaviors is solved, and the recognition efficiency and accuracy are improved.

CN114494961BActive Publication Date: 2025-11-18GUANGZHOU FUGANG LIFE INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210081754.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-24
Publication Date
2025-11-18
Estimated Expiration
2042-01-24

AI Technical Summary

Technical Problem

In existing technologies, multi-category abnormal behavior recognition requires multiple recognition models to be trained and loaded separately, which is complex and makes it difficult to effectively reduce the recognition complexity.

Method used

The sample videos are grouped and fused to generate fused sample videos, which are used to train the recognition model to identify multiple categories of abnormal behaviors.

Benefits of technology

It improves the training efficiency of the recognition model and the recognition efficiency of multi-class abnormal behaviors, reduces the recognition complexity, and achieves efficient recognition of multi-class abnormal behaviors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114494961B_ABST
    Figure CN114494961B_ABST
Patent Text Reader

Abstract

The application discloses a sample data processing method and device applied to model training, comprising the following steps: obtaining a plurality of sample videos to be trained; grouping the plurality of sample videos according to the categories of all abnormal behaviors to be recognized to obtain a plurality of sample video groups; performing a fusion operation on each sample video group to obtain a fusion sample video corresponding to each sample video group; and training a pre-determined recognition model according to all the fusion sample videos to obtain a trained recognition model, which is used to identify whether there is an abnormal behavior in any target video and further determine the specific abnormal behavior involved when there is. It can be seen that the application can train the recognition model according to the fusion sample video generated by the plurality of sample videos, which can not only improve the training efficiency of the recognition model, but also realize the identification of multi-category abnormal behaviors, reduce the identification complexity of multi-category abnormal behaviors and improve the identification efficiency of multi-category abnormal behaviors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of information processing technology, and in particular to a method and apparatus for processing sample data used in model training. Background Technology

[0002] With the rapid development of electronic technology, in the field of intelligent monitoring technology, in addition to collecting video to achieve real-time monitoring of the monitored area, it is also necessary to further analyze the video content contained in the collected video in order to identify whether there are abnormal behaviors in the video, such as people falling or local high temperature behavior.

[0003] In practical applications, to identify abnormal behavior in captured videos, besides manual identification, a pre-determined recognition model can be trained using a large number of sample videos. This trained model can then be applied to intelligent monitoring technology to achieve real-time identification of abnormal behavior in the captured videos and further enable intelligent alerts for abnormal behavior, allowing relevant personnel to quickly and effectively handle such situations. However, practice has shown that the abnormal behavior categories in the sample videos currently used to train the recognition model are relatively limited. To identify multiple categories of abnormal behavior, it is necessary to train the recognition model separately using multiple sample video sets, resulting in multiple recognition models for different categories, each used to identify different types of abnormal behavior. Furthermore, when identifying multiple categories of abnormal behavior, multiple recognition models of different categories need to be loaded, making the operation complex. Therefore, reducing the complexity of identifying multiple categories of abnormal behavior is particularly important. Summary of the Invention

[0004] This invention provides a sample data processing method and apparatus for model training, which can reduce the complexity of identifying multiple categories of abnormal behaviors.

[0005] The first aspect of this invention discloses a sample data processing method for model training, the method comprising:

[0006] Obtain several sample videos to be used for training;

[0007] Based on the abnormal behavior categories of all target abnormal behaviors to be identified, the sample videos are grouped to obtain several sample video groups. Each sample video group includes at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group include all target abnormal behaviors.

[0008] Perform a video fusion operation on each of the sample video groups to obtain a fused sample video corresponding to each sample video group;

[0009] The pre-determined identification model is trained according to all the fusion sample videos, and a trained identification model is obtained, which is used to identify whether there is an abnormal behavior in the video content of any target video and determine at least one target abnormal behavior existing in the video content of the target video when it is identified that there is an abnormal behavior in the video content of the target video.

[0010] As an optional implementation, in the first aspect of the present application, the performing a video fusion operation on each of the sample video groups to obtain a corresponding fusion sample video of each of the sample video groups comprises:

[0011] For each of the sample video groups, a first video splicing operation is performed on all sample videos in the sample video group to obtain a corresponding fusion sample video of the sample video group; wherein the first video splicing operation is a time vertical splicing operation, and the video content of the fusion sample video of the sample video group at any time of its corresponding video duration includes the video content of each sample video in the sample video group at the corresponding time.

[0012] As an optional implementation, in the first aspect of the present application, before the performing a first video splicing operation on all sample videos in the sample video group to obtain a corresponding fusion sample video of the sample video group, the method further comprises:

[0013] For each of the sample video groups, the duration of each sample video in the sample video group is counted and the sum of the durations of all sample videos in the sample video group is calculated to obtain a total video duration corresponding to the sample video group; it is judged whether the total video duration corresponding to the sample video group is less than or equal to a pre-determined video duration threshold, when the result of the judgment is no, the performing a first video splicing operation on all sample videos in the sample video group to obtain a corresponding fusion sample video of the sample video group is triggered; when the result of the judgment is yes, a second video splicing operation is performed on all sample videos in the sample video group to obtain a corresponding fusion sample video of the sample video group, wherein the second video splicing operation is a time horizontal splicing operation, and the video content of the fusion sample video of the sample video group at any time of its corresponding video duration is the video content of one of the sample videos in the sample video group at the corresponding time.

[0014] As an optional implementation, in the first aspect of the present application, after the obtaining a plurality of sample videos to be trained, the method further comprises:

[0015] It is judged whether there is at least one target sample video in all the sample videos, wherein the abnormal behavior annotation information corresponding to the target sample video includes two or more than two abnormal behavior categories.

[0016] When it is judged that there is at least one of the target sample video, the training abnormal behavior category corresponding to each of the target sample video is screened from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each of the target sample video according to the predetermined abnormal behavior screening factor, and all abnormal behavior categories except the screened training abnormal behavior category in the abnormal behavior annotation information corresponding to each of the target sample video are hidden, so as to update the abnormal behavior annotation information corresponding to all of the target sample video.

[0017] As an optional implementation, in the first aspect of the present application, the grouping of the plurality of sample videos according to the abnormal behavior categories of all target abnormal behaviors to be identified to obtain a plurality of sample video groups comprises:

[0018] According to the abnormal behavior categories of all target abnormal behaviors to be identified and the abnormal behavior annotation information corresponding to each of the sample videos, all of the sample videos are divided into a plurality of sample video sets, each of the sample video sets comprises a plurality of sample videos, one of the sample video sets corresponds to one of the abnormal behavior categories of the target abnormal behaviors, the abnormal behavior annotation information corresponding to the sample videos in the same sample video set is the same, and the abnormal behavior annotation information corresponding to the sample videos in different sample video sets is different;

[0019] One sample video is selected from each of the sample video sets respectively, and the sample video selected from each of the sample video sets respectively is determined as a sample video group;

[0020] It is judged whether the number of all the determined sample video groups reaches a first preset number threshold, when the judgment result is no, the operation of selecting one sample video from each of the sample video sets respectively and determining the sample video selected from each of the sample video sets respectively as a sample video group is repeatedly executed until the number of all the determined sample video groups reaches the first preset number threshold.

[0021] As an optional implementation, in the first aspect of the present application, before the operation of selecting one sample video from each of the sample video sets respectively and determining the sample video selected from each of the sample video sets respectively as a sample video group, the method further comprises:

[0022] For any one of the sample video sets, it is judged whether the total number of sample videos included in the sample video set is less than a second preset number threshold; when it is judged that the total number of sample videos included in the sample video set is less than the second preset number threshold, a sample video supplement operation of the same abnormal behavior category is performed on the sample video set according to the abnormal behavior category corresponding to the sample video set until the total number of sample videos included in the sample video set is not less than the second preset number threshold.

[0023] As an optional implementation, in the first aspect of the application, the selecting one sample video from each of the sample video sets respectively comprises:

[0024] selecting one sample video from all the sample videos in one of the sample video sets which have not been selected as the reference sample video, obtaining the video information corresponding to the reference sample video as the reference information, and respectively screening sample videos whose matching degrees of corresponding video information with the reference information are greater than or equal to a preset matching degree threshold from all the sample videos in the remaining sample video sets which have not been selected.

[0025] The second aspect of the application discloses a sample data processing device applied to model training, the device comprises:

[0026] a sample acquisition module, configured to acquire a plurality of sample videos to be trained;

[0027] a sample grouping module, configured to group the plurality of sample videos according to the abnormal behavior categories of all target abnormal behaviors to be recognized, to obtain a plurality of sample video groups, each of the sample video groups comprising at least two sample videos, and all abnormal behaviors involved in all sample videos included in each of the sample video groups comprising all target abnormal behaviors;

[0028] a sample fusion module, configured to perform a video fusion operation on each of the sample video groups to obtain a fusion sample video corresponding to each of the sample video groups;

[0029] a model training module, configured to train a pre-determined recognition model according to all fusion sample videos to obtain a trained recognition model, the trained recognition model being used to identify whether there is an abnormal behavior in the video content of any target video and to determine at least one target abnormal behavior existing in the video content of the target video when it is identified that there is an abnormal behavior in the video content of the target video.

[0030] As an optional implementation, in the second aspect of the application, the sample fusion module comprises:

[0031] a vertical splicing unit, configured to perform a first video splicing operation on all sample videos in each of the sample video groups to obtain a fusion sample video corresponding to the sample video group, wherein the first video splicing operation is a time vertical splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any time within the video duration corresponding to the fusion sample video comprises the video content of each sample video in the sample video group at the corresponding time.

[0032] As an optional implementation form, in the second aspect, the sample fusion module further comprises:

[0033] a duration judging unit, configured to, before the vertical splicing unit performs the first video splicing operation on all sample videos in each of the sample video groups to obtain a fusion sample video corresponding to the sample video group, count the duration of each sample video in the sample video group and calculate the sum of the durations of all sample videos in the sample video group to obtain a total video duration corresponding to the sample video group, and judge whether the total video duration corresponding to the sample video group is less than or equal to a predetermined video duration threshold, and when it is judged that the total video duration corresponding to the sample video group is not less than or equal to the predetermined video duration threshold, trigger the vertical splicing unit to perform the operation of performing the first video splicing operation on all sample videos in the sample video group to obtain a fusion sample video corresponding to the sample video group.

[0034] a horizontal splicing unit, configured to, when the duration judging unit judges that the total video duration corresponding to each of the sample video groups is greater than the video duration threshold, perform a second video splicing operation on all sample videos in the sample video group to obtain a fusion sample video corresponding to the sample video group, wherein the second video splicing operation is a time horizontal splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any time within the video duration corresponding to the fusion sample video is the video content of one of the sample videos in the sample video group at the corresponding time.

[0035] As an optional implementation form, in the second aspect, the device further comprises:

[0036] a label judging module, configured to, after the sample obtaining module obtains the plurality of sample videos to be trained, judge whether there is at least one target sample video in all the sample videos, wherein the abnormal behavior label information corresponding to the target sample video comprises two or more than two abnormal behavior categories.

[0037] The label updating module is configured to, when the label judging module judges that there is at least one target sample video, screen a training abnormal behavior category corresponding to each target sample video from all abnormal behavior categories included in abnormal behavior label information corresponding to each target sample video according to a predetermined abnormal behavior screening factor, and hide all abnormal behavior categories except the screened training abnormal behavior category from the abnormal behavior label information corresponding to each target sample video, so as to update the abnormal behavior label information corresponding to all target sample videos.

[0038] As an optional implementation, in the second aspect of the present application, the sample grouping module comprises:

[0039] The sample dividing unit is configured to divide all sample videos into a plurality of sample video sets according to abnormal behavior categories of all target abnormal behaviors to be recognized and abnormal behavior label information corresponding to each sample video, each sample video set comprising a plurality of sample videos, one sample video set corresponding to one abnormal behavior category of a target abnormal behavior, and the abnormal behavior label information corresponding to sample videos in the same sample video set being the same, and the abnormal behavior label information corresponding to sample videos in different sample video sets being different.

[0040] The sample determining unit is configured to select one sample video from each sample video set respectively, and determine the sample video selected from each sample video set respectively as one sample video group.

[0041] The number judging unit is configured to judge whether the number of all sample video groups determined reaches a first preset number threshold, and when the judgment result is no, trigger the sample determining unit to perform the operation of selecting one sample video from each sample video set respectively, and determining the sample video selected from each sample video set respectively as one sample video group, until the number of all sample video groups determined reaches the first preset number threshold.

[0042] As an optional implementation, in the second aspect of the present application, the sample judging unit is further configured to, before the sample determining unit selects one sample video from each sample video set respectively, and determines the sample video selected from each sample video set respectively as one sample video group, judge, for any sample video set, whether the total number of sample videos included in the sample video set is less than a second preset number threshold.

[0043] The sample grouping module further comprises:

[0044] The sample supplement unit is configured to, when the sample judgment module judges that the total number of sample videos included in the sample video set is less than the second preset number threshold, perform a sample video supplement operation of the same abnormal behavior category on the sample video set according to the abnormal behavior category corresponding to the sample video set until the total number of sample videos included in the sample video set is not less than the second preset number threshold.

[0045] As an optional implementation, in the second aspect of the present application, the specific manner of selecting one sample video from each sample video set by the sample determination unit comprises:

[0046] selecting one sample video from all the sample videos in one of the sample video sets that have not been selected as a reference sample video, obtaining the video information corresponding to the reference sample video as reference information, and respectively screening sample videos from all the sample videos in the remaining sample video sets that have not been selected, the matching degree of the video information corresponding to the sample videos being greater than or equal to a preset matching degree threshold.

[0047] The third aspect of the present application discloses another sample data processing device applied to model training, the device comprises:

[0048] a memory storing executable program codes;

[0049] a processor coupled with the memory;

[0050] The processor invokes the executable program codes stored in the memory to execute part or all of the steps of the sample data processing method applied to model training disclosed in the first aspect of the present application.

[0051] The fourth aspect of the present application discloses a computer storage medium, the computer storage medium stores computer instructions, when the computer instructions are invoked, part or all of the steps of the sample data processing method applied to model training disclosed in the first aspect of the present application are executed.

[0052] Compared with the prior art, the present application has the following beneficial effects:

[0053] Obtain a plurality of sample videos to be trained; group the plurality of sample videos according to all target abnormal behaviors of the abnormal behavior categories to be identified, to obtain a plurality of sample video groups, each sample video group including at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group including all target abnormal behaviors; perform a video fusion operation on each sample video group to obtain a fusion sample video corresponding to each sample video group; and train a pre-determined identification model according to all fusion sample videos to obtain a trained identification model, the trained identification model being used to identify whether there is an abnormal behavior in the video content of any target video and to determine at least one target abnormal behavior existing in the video content of the target video when identifying that there is an abnormal behavior in the video content of the target video. It can be seen that the present application can train the identification model according to the fusion sample video generated by the plurality of sample videos, which not only can improve the training efficiency of the identification model, but also can realize the identification of multi-category abnormal behaviors, reduce the identification complexity of multi-category abnormal behaviors, and improve the identification efficiency of multi-category abnormal behaviors. BRIEF DESCRIPTION OF DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0055] Figure 1 is a flowchart of a sample data processing method applied to model training disclosed by the embodiments of the present application;

[0056] Figure 2 is a flowchart of another sample data processing method applied to model training disclosed by the embodiments of the present application;

[0057] Figure 3 is a structural diagram of a sample data processing device applied to model training disclosed by the embodiments of the present application;

[0058] Figure 4 is a structural diagram of another sample data processing device applied to model training disclosed by the embodiments of the present application;

[0059] Figure 5 is a structural diagram of another sample data processing device applied to model training disclosed by the embodiments of the present application. DETAILED DESCRIPTION

[0060] In the following, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0061] The terms "first", "second", and the like in the specification and claims of the present application and the above-described drawings are used to distinguish different objects, rather than to describe a particular order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product, port, or terminal that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed or can optionally include other steps or units inherent to the process, method, product, port, or terminal.

[0062] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearance of the phrase in various places in the specification does not necessarily all refer to the same embodiment, nor does it necessarily refer to a separate or alternative embodiment in isolation from other embodiments. It is explicitly and implicitly understood by those skilled in the art that embodiments described herein can be combined with other embodiments.

[0063] The present application discloses a sample data processing method and device applied to model training, which can train the recognition model according to the fusion sample video generated by a plurality of sample videos, can not only improve the training efficiency of the recognition model, but also realize the recognition of multi-class abnormal behaviors, reduce the recognition complexity of multi-class abnormal behaviors, and improve the recognition efficiency of multi-class abnormal behaviors. The following will be described in detail.

[0064] Embodiment one

[0065] Please refer to Figure 1 , Figure 1 is a flowchart of a sample data processing method applied to model training disclosed by the embodiments of the present application. Wherein, Figure 1 The described method can be applied to a sample data processing device. As Figure 1 shown, the sample data processing method applied to model training can include the following operations:

[0066] 101, obtaining a plurality of sample videos to be trained.

[0067] In the embodiment of the present application, each sample video has corresponding abnormal behavior annotation information, and the abnormal behavior annotation information of the sample video is used to specifically annotate the abnormal behavior category included in the video content of the sample video. If the abnormal behavior annotation information is empty, it means that there is no abnormal behavior to be identified in the video content of the sample video. Further, when the video content of the sample video has abnormal behavior, the abnormal behavior annotation information corresponding to the sample video includes at least one abnormal behavior category.

[0068] Further, one abnormal behavior category can correspond to at least one abnormal behavior. Preferably, one abnormal behavior category corresponds to one abnormal behavior, which is beneficial to improve the accuracy of the trained identification model in identifying specific abnormal behaviors.

[0069] 102. Group a plurality of sample videos according to the abnormal behavior categories of all target abnormal behaviors to be identified to obtain a plurality of sample video groups.

[0070] In the embodiment of the present application, each sample video group includes at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group include all target abnormal behaviors. Among them, for a certain target abnormal behavior, there is at least one sample video in each sample video group, and the corresponding abnormal behavior annotation information of the sample video includes the abnormal behavior category corresponding to the target abnormal behavior. Preferably, the abnormal behavior annotation information corresponding to different sample videos in the sample video group includes different abnormal behavior categories.

[0071] In the embodiment of the present application, the grouping of a plurality of sample videos according to the abnormal behavior categories of all target abnormal behaviors to be identified improves the accuracy of the sample video group and the comprehensiveness of all abnormal behaviors involved.

[0072] 103. Perform a video fusion operation on each sample video group to obtain a fusion sample video corresponding to each sample video group.

[0073] 104. Train a pre-determined identification model according to all fusion sample videos to obtain a trained identification model.

[0074] It can be seen that the method described in the embodiment of the present application can train the identification model according to the fusion sample video generated by a plurality of sample videos, which not only improves the training efficiency of the identification model, but also realizes the identification of multi-category abnormal behaviors, reduces the identification complexity of multi-category abnormal behaviors, and improves the identification efficiency of multi-category abnormal behaviors.

[0075] In an optional embodiment, performing a video fusion operation on each sample video group to obtain a fusion sample video corresponding to each sample video group can include:

[0076] For each sample video group, a first video splicing operation is performed on all sample videos in the sample video group to obtain a fusion sample video corresponding to the sample video group; wherein the first video splicing operation is a time vertical splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any moment of the video duration thereof includes the video content of each sample video in the sample video group at the corresponding moment.

[0077] Optionally, when the video durations of different sample videos in the sample video group are different and the difference is large, a sample video containing at least a video segment corresponding to the abnormal behavior can be obtained by performing a clipping operation on the sample video. Further optionally, the sample video containing the video segment corresponding to the abnormal behavior can also include a video segment corresponding to the cause of the abnormal behavior and / or a video segment corresponding to the consequence of the abnormal behavior. In this way, it is beneficial to enable the trained recognition model to not only recognize the abnormal behavior, but also recognize the cause of the abnormal behavior and the consequence of the abnormal behavior, wherein the cause of the abnormal behavior helps to realize early warning of the abnormal behavior to reduce the probability of occurrence of the abnormal behavior, and the consequence of the abnormal behavior helps to realize the processing guidance of the abnormal behavior.

[0078] It can be seen that the optional embodiment can perform a time vertical splicing operation on the sample videos in the sample video group, which is beneficial to reduce the video duration of the fusion sample video while realizing the diversity of the abnormal behavior or the abnormal behavior category involved in the fusion sample video, and is beneficial to improve the training efficiency of training the recognition model based on the fusion sample video.

[0079] In the optional embodiment, further optionally, before performing the first video splicing operation on all sample videos in the sample video group to obtain the fusion sample video corresponding to the sample video group, the method can further include the following operations:

[0080] For each sample video group, the duration of each sample video in the sample video group is counted and the sum of the durations of all sample videos in the sample video group is calculated to obtain the total video duration corresponding to the sample video group; it is judged whether the total video duration corresponding to the sample video group is less than or equal to a predetermined video duration threshold, and when the judgment result is no, the operation of performing the first video splicing operation on all sample videos in the sample video group to obtain the fusion sample video corresponding to the sample video group is triggered. Further optionally, when the judgment result is yes, a second video splicing operation is performed on all sample videos in the sample video group to obtain the fusion sample video corresponding to the sample video group, wherein the second video splicing operation is a time horizontal splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any moment of the video duration thereof is the video content of one of the sample videos in the sample video group at the corresponding moment.

[0081] It can be seen that the optional embodiment can also intelligently judge the total video duration of the sample videos in the sample video set before performing the vertical splicing operation on the sample videos, and then perform the matched splicing operation, thereby providing diversified video splicing modes, and ensuring the comprehensiveness of the abnormal behavior categories involved in the fused sample videos while controlling the video duration, which is beneficial to improve the training effect and training efficiency of the subsequent recognition model.

[0082] Embodiment two

[0083] Please refer to Figure 2 , Figure 2 is another flowchart of the sample data processing method applied to model training disclosed by the embodiment of the application. Wherein, Figure 2 The method described can be applied to a sample data processing device. As Figure 2 shown, the sample data processing method applied to model training can include the following operations:

[0084] 201, obtaining a plurality of sample videos to be trained.

[0085] 202, judging whether there is at least one target sample video in all sample videos, the abnormal behavior annotation information corresponding to the target sample video including two or more than two abnormal behavior categories, when the judgment result of step 202 is yes, triggering step 203- step 204; when the judgment result of step 202 is no, step 205 can be triggered.

[0086] 203, according to the pre-determined abnormal behavior screening factor, screening the training abnormal behavior category corresponding to each target sample video from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each target sample video.

[0087] Optionally, the abnormal behavior screening factor can be the significant level of abnormal behavior, and the abnormal behavior category corresponding to the abnormal behavior with higher significant level is screened as the training abnormal behavior category from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each target sample video. Wherein, the significant level of abnormal behavior indicates that the more obvious the abnormal behavior contained in the video content of the sample video. In this way, it is beneficial to reduce the influence of multi-annotation information on the subsequent training process, which is not only beneficial to improve the accuracy of the sample video, but also beneficial to improve the training efficiency of the subsequent recognition model.

[0088] In other optional embodiments, the abnormal behavior screening factor can specifically be the number of sample videos corresponding to the abnormal behavior category. Prioritize the selection of the abnormal behavior category with the fewest corresponding sample videos from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each target sample video as the abnormal behavior category for training. This is beneficial for the rapid expansion of missing sample videos and reduces the occurrence of unreasonable utilization of sample video resources.

[0089] 204. Hide all abnormal behavior categories in the abnormal behavior annotation information corresponding to each target sample video except for the selected abnormal behavior categories used for training, in order to update the abnormal behavior annotation information corresponding to all target sample videos.

[0090] In an optional embodiment, for any target sample video, before hiding a certain anomalous behavior category other than the selected training anomalous behavior categories, it is determined whether the number of sample videos corresponding to that anomalous behavior category is less than or equal to the required number. If the determination result is yes, a copying operation is performed on the target sample video to obtain a copied sample video, and the certain anomalous behavior category is determined as the training anomalous behavior category corresponding to the copied sample video. Then, the certain anomalous behavior category of the target sample video is hidden, and the copied sample video is used as the sample video corresponding to that anomalous behavior category. This facilitates the rapid expansion of missing sample videos, promotes the rational use of sample resources, and improves the utilization rate of sample resources.

[0091] 205. Based on the abnormal behavior category of all target abnormal behaviors to be identified, several sample videos are grouped to obtain several sample video groups.

[0092] In this embodiment of the invention, each sample video group includes at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group include all target abnormal behaviors.

[0093] 206. Perform video fusion operation on each sample video group to obtain the fused sample video corresponding to each sample video group.

[0094] 207. Train the pre-determined recognition model based on all fused sample videos to obtain the trained recognition model.

[0095] In this embodiment of the invention, the relevant descriptions of steps 205-207 are as described in the detailed descriptions of steps 102-104 in Embodiment 1, and will not be repeated in this embodiment of the invention.

[0096] It can be seen that the method described in the embodiment of the present application can train the recognition model according to the fusion sample video generated by the plurality of sample videos, which can not only improve the training efficiency of the recognition model, but also realize the recognition of multi-class abnormal behaviors, reduce the recognition complexity of multi-class abnormal behaviors, and improve the recognition efficiency of multi-class abnormal behaviors. In addition, it is beneficial to reduce the influence of multi-label information on the subsequent training process, which is not only beneficial to improve the accuracy of the sample video, but also beneficial to improve the training efficiency of the subsequent recognition model. In addition, it is beneficial to realize the rapid expansion of the missing sample video, and is beneficial to realize the rational use of sample resources and improve the utilization rate of sample resources.

[0097] In an optional embodiment, according to the predetermined abnormal behavior screening factor, the training abnormal behavior category corresponding to each target sample video is screened from all abnormal behavior categories included in the abnormal behavior label information corresponding to each target sample video, which can include:

[0098] According to the abnormal behavior category of all target abnormal behaviors to be recognized and the abnormal behavior label information corresponding to each sample video, all sample videos are divided into a plurality of sample video sets, each sample video set includes a plurality of sample videos, and one sample video set corresponds to an abnormal behavior category of a target abnormal behavior. The abnormal behavior label information corresponding to the sample videos in the same sample video set is the same, and the abnormal behavior label information corresponding to the sample videos in different sample video sets is different;

[0099] Respectively selecting one sample video from each sample video set, and determining the sample video selected from each sample video set as a sample video group;

[0100] Determine whether the number of all determined sample video groups reaches a first preset number threshold, and when the determination result is no, repeat the operation of respectively selecting one sample video from each sample video set, and determining the sample video selected from each sample video set as a sample video group until the number of all determined sample video groups reaches the first preset number threshold.

[0101] It can be seen that the optional embodiment can also divide sample videos of the same abnormal behavior category into the same sample video set through grouping, and then select sample videos for determining a sample video group from each sample video set, which is beneficial to improve the efficiency, accuracy and matching degree with the actual training requirements of the recognition model of the determined sample video group.

[0102] In the optional embodiment, further optionally, before respectively selecting one sample video from each sample video set and determining the sample video selected from each sample video set as a sample video group, the method can further include the following operations:

[0103] For any sample video set, it is judged whether the total number of sample videos included in the sample video set is less than a second preset number threshold. When it is judged that the total number of sample videos included in the sample video set is less than the second preset number threshold, a sample video supplement operation of the same abnormal behavior category is performed on the sample video set according to the abnormal behavior category corresponding to the sample video set, until the total number of sample videos included in the sample video set is not less than the second preset number threshold.

[0104] Optionally, the sample video supplement operation of the same abnormal behavior category performed on the sample video set according to the abnormal behavior category corresponding to the sample video set can include:

[0105] An operation matched with the preset condition is performed on the sample video in the sample video set that meets the preset condition, to obtain the sample video of the same abnormal behavior category.

[0106] Further optionally, the preset condition can be that the sample video includes multiple video segments of the same abnormal behavior, and the operation matched with the preset condition can be a sample video division operation, which is used to divide the sample video into multiple sub-videos, each of which has the abnormal behavior and all the sub-videos divided from the sample video have the same abnormal behavior category, and each sub-video is determined as the sample video of the same abnormal behavior category and added to the corresponding sample video set. Alternatively, the preset condition can also be that a new sample video is generated by performing a processing operation on the sample video, and the new sample video has the same abnormal behavior category as the original sample video. Optionally, the processing operation can include at least one of an environmental parameter transformation operation of video content, a filter parameter transformation operation of video content, a relocation operation of an abnormal behavior corresponding object, etc.

[0107] It can be seen that the optional embodiment can further judge the number of sample videos in each sample video set, and then supplement the sample videos of the same abnormal category in the sample video set in time when the number of sample videos is insufficient, which is beneficial to improve the accuracy and efficiency of the subsequently determined sample video group.

[0108] In the optional embodiment, further optionally, selecting one sample video from each sample video set can include:

[0109] select one of the sample videos from all the sample videos in the one of the sample video sets that have not been selected as the reference sample video, obtain video information corresponding to the reference sample video as reference information, and respectively filter sample videos whose corresponding video information has a matching degree greater than or equal to a preset matching degree threshold from all the sample videos in the remaining sample video sets that have not been selected.

[0110] Optionally, the video information corresponding to the reference sample video can include one or a combination of the following: a video source, a video duration, a video canvas size, a time period or time point at which an abnormal behavior appears in the video, an environmental parameter in the video, an object parameter of an object corresponding to the abnormal behavior in the video, and a trigger factor corresponding to the abnormal behavior in the video.

[0111] It can be seen that the optional embodiment can also filter sample videos whose corresponding video information has a matching degree greater than or equal to a preset matching degree threshold to obtain a sample video group, which is beneficial to improving the consistency between the sample videos in the determined sample video group, and is also beneficial to improving the training efficiency and training accuracy.

[0112] Embodiment Three

[0113] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of a sample data processing device applied to model training disclosed by an embodiment of the present application. As Figure 3 shown, the sample data processing device applied to model training can include:

[0114] The sample acquisition module 301 is configured to acquire a plurality of sample videos to be trained.

[0115] The sample grouping module 302 is configured to group the plurality of sample videos according to abnormal behavior categories of all target abnormal behaviors to be recognized, to obtain a plurality of sample video groups, each sample video group including at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group including all target abnormal behaviors.

[0116] The sample fusion module 303 is configured to perform a video fusion operation on each sample video group to obtain a fusion sample video corresponding to each sample video group.

[0117] The model training module 304 is configured to train a pre-determined recognition model according to all fusion sample videos to obtain a trained recognition model, the trained recognition model being used to identify whether there is an abnormal behavior in video content of any target video and to determine at least one target abnormal behavior existing in the video content of the target video when it is identified that there is an abnormal behavior in the video content of the target video.

[0118] It can be seen that the embodimentFigure 3 The described device can train the recognition model according to the fusion sample video pairs generated by the plurality of sample videos, can not only improve the training efficiency of the recognition model, but also realize the recognition of multi-class abnormal behaviors, reduce the recognition complexity of the multi-class abnormal behaviors, and improve the recognition efficiency of the multi-class abnormal behaviors.

[0119] In an optional embodiment, as shown in Figure 4 The sample fusion module 303 can include, as shown in

[0120] The vertical splicing unit 3031 is configured to, for each sample video group, perform a first video splicing operation on all sample videos in the sample video group to obtain a fusion sample video corresponding to the sample video group; wherein the first video splicing operation is a time vertical splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any time within the video duration corresponding to the fusion sample video includes the video content of each sample video in the sample video group at the corresponding time.

[0121] In this optional embodiment, the sample fusion module 303 can further include, further optionally,

[0122] The duration judgment unit 3032 is configured to, for each sample video group, before the vertical splicing unit 3031 performs the first video splicing operation on all sample videos in the sample video group to obtain the fusion sample video corresponding to the sample video group, count the duration of each sample video in the sample video group and calculate the sum of the durations of all sample videos in the sample video group to obtain the total video duration corresponding to the sample video group; determine whether the total video duration corresponding to the sample video group is less than or equal to a predetermined video duration threshold, and when it is determined that the total video duration corresponding to the sample video group is not less than or equal to the predetermined video duration threshold, trigger the vertical splicing unit 3031 to perform the operation of performing the first video splicing operation on all sample videos in the sample video group to obtain the fusion sample video corresponding to the sample video group.

[0123] The horizontal splicing unit 3033 is configured to, for each sample video group, when the duration judgment unit 3032 determines that the total video duration corresponding to the sample video group is greater than the video duration threshold, perform a second video splicing operation on all sample videos in the sample video group to obtain a fusion sample video corresponding to the sample video group, wherein the second video splicing operation is a time horizontal splicing operation, and the video content of the fusion sample video corresponding to the sample video group at any time within the video duration corresponding to the fusion sample video is the video content of one of the sample videos in the sample video group at the corresponding time.

[0124] It can be seen that the optional embodiment can also perform the time vertical splicing operation on the sample videos in the sample video group, which is beneficial to reduce the video length of the fusion sample video while achieving the diversity of the abnormal behaviors or the abnormal behavior categories involved in the fusion sample video, and is beneficial to improve the training efficiency of the subsequent training of the identification model based on the fusion sample video. In addition, the total video length of the sample videos in the sample video group can be intelligently judged before the vertical splicing operation is performed, and then the matching splicing operation is performed, which not only provides diversified video splicing methods, but also ensures the comprehensiveness of the abnormal behavior categories involved in the fusion sample video while controlling the video length, which is beneficial to improve the training effect and training efficiency of the subsequent training of the identification model.

[0125] In another optional embodiment, as shown in Figure 4 The device can further include:

[0126] The labeling judgment module 305 is configured to, after the sample acquisition module 301 acquires the plurality of sample videos to be trained, judge whether there is at least one target sample video in all the sample videos, wherein the abnormal behavior labeling information corresponding to the target sample video includes two or more abnormal behavior categories.

[0127] The labeling update module 306 is configured to, when the labeling judgment module 305 judges that there is at least one target sample video, screen the training abnormal behavior category corresponding to each target sample video from all the abnormal behavior categories included in the abnormal behavior labeling information corresponding to each target sample video according to the pre-determined abnormal behavior screening factor, and hide all the abnormal behavior categories in the abnormal behavior labeling information corresponding to each target sample video except the screened training abnormal behavior category, to update the abnormal behavior labeling information corresponding to all the target sample videos.

[0128] It should be noted that when the labeling judgment module 305 judges that there is no target sample video, the sample grouping module 302 can directly trigger the above-mentioned operation of grouping the plurality of sample videos according to the abnormal behavior categories of all the target abnormal behaviors to be identified to obtain the plurality of sample video groups.

[0129] It can be seen that the optional embodiment can also be beneficial to reduce the influence of multi-labeling information on the subsequent training process, which is not only beneficial to improve the accuracy of the sample video, but also beneficial to improve the training efficiency of the subsequent identification model.

[0130] In yet another optional embodiment, the sample grouping module 302 can include:

[0131] The sample division unit 3021 is configured to divide all the sample videos into a plurality of sample video sets according to the abnormal behavior categories of all the target abnormal behaviors to be recognized and the abnormal behavior annotation information corresponding to each sample video, each sample video set including a plurality of sample videos, one sample video set corresponding to one abnormal behavior category of a target abnormal behavior, the sample videos in the same sample video set having the same abnormal behavior annotation information, and the sample videos in different sample video sets having different abnormal behavior annotation information.

[0132] The sample determination unit 3022 is configured to select one sample video from each sample video set respectively, and determine the sample videos selected from each sample video set respectively as one sample video group.

[0133] The number judgment unit 3023 is configured to judge whether the number of all the sample video groups determined reaches a first preset number threshold, and when the judgment result is no, trigger the sample determination unit 3022 to perform the operation of selecting one sample video from each sample video set respectively and determining the sample videos selected from each sample video set respectively as one sample video group until the number of all the sample video groups determined reaches the first preset number threshold.

[0134] In the optional embodiment, further optionally, the sample judgment unit 3023 is further configured to, before the sample determination unit 3022 selects one sample video from each sample video set respectively and determines the sample videos selected from each sample video set respectively as one sample video group, judge, for any sample video set, whether the total number of sample videos included in the sample video set is less than a second preset number threshold. Wherein, as shown in the figure, Figure 4 The sample grouping module 302 can further include:

[0135] The sample supplement unit 3024 is configured to, when the sample judgment module 3023 judges that the total number of sample videos included in the sample video set is less than the second preset number threshold, perform a sample video supplement operation of the same abnormal behavior category on the sample video set according to the abnormal behavior category corresponding to the sample video set until the total number of sample videos included in the sample video set is not less than the second preset number threshold.

[0136] It can be seen that the optional embodiment can also divide the sample videos of the same abnormal behavior category into the same sample video set by grouping, and then screen the sample videos for determining the sample video group from each sample video set, which is beneficial to improve the efficiency, accuracy and matching degree with the actual training requirement of the identified model of the determined sample video group. In addition, the number of sample videos in each sample video set is judged, and then the sample videos of the same abnormal category in the sample video set are supplemented in time when the number of sample videos is insufficient, which is beneficial to improve the accuracy and efficiency of the subsequently determined sample video group.

[0137] In the optional embodiment, further optionally, the specific manner of the sample determination unit 3022 selecting one sample video from each sample video set can include:

[0138] selecting one sample video from all the sample videos in one sample video set that have not been selected as a reference sample video, obtaining the video information corresponding to the reference sample video as reference information, and screening sample videos with a matching degree of corresponding video information greater than or equal to a preset matching degree threshold from all the sample videos in the remaining sample video sets that have not been selected.

[0139] It can be seen that the optional embodiment can also screen sample videos with a matching degree of corresponding video information greater than or equal to a preset matching degree threshold to obtain a sample video group, which is beneficial to improve the consistency between the sample videos in the determined sample video group, and also beneficial to improve the training efficiency and training accuracy.

[0140] Embodiment Four

[0141] Please refer to Figure 5 , Figure 5 is another structure diagram of a sample data processing device for model training disclosed by the embodiments of the present application. As shown in Figure 5 , the sample data processing device for model training can include:

[0142] a memory 401 storing executable program codes;

[0143] a processor 402 coupled with the memory 401;

[0144] The processor 402 calls the executable program codes stored in the memory 401 to execute part or all of the steps of the sample data processing method for model training disclosed in the embodiments one or two of the present application.

[0145] Embodiment Five

[0146] The embodiment of the present application discloses a computer storage medium, the computer storage medium stores computer instructions, when the computer instructions are invoked, part or all steps of the sample data processing method applied to model training disclosed in the embodiment one or the embodiment two are executed.

[0147] The device embodiments described above are only schematic, wherein the modules described as separate components can or can not be physically separated, and the components displayed as modules can or can not be physical modules, i.e., can be located in one place or distributed on multiple network modules. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0148] Through the specific description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be realized by means of software and the necessary general hardware platform, and of course, can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, the storage medium includes a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a programmable read-only memory (Programmable Read-only Memory, PROM), an erasable programmable read-only memory (Erasable Programmable Read Only Memory, EPROM), a one-time programmable read-only memory (One-time Programmable Read-Only Memory, OTPROM), an electrically erasable programmable read-only memory (Electrically-Erasable Programmable Read-Only Memory, EEPROM), a compact disc read-only memory (Compact Disc Read-Only Memory, CD-ROM) or other optical disk storage, a magnetic disk storage, a magnetic tape storage, or any other computer readable medium that can be used to carry or store data.

[0149] It should be finally pointed out that: the sample data processing method and device applied to model training disclosed by the embodiments of the present application are only the preferred embodiments of the present application, and are used to illustrate the technical solutions of the present application, but not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that; it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A sample data processing method for model training, characterized in that, The method includes: Obtain several sample videos to be used for training; Based on the abnormal behavior categories of all target abnormal behaviors to be identified, the sample videos are grouped to obtain several sample video groups. Each sample video group includes at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group include all target abnormal behaviors. For each of the sample video groups, a first video splicing operation is performed on all the sample videos in the sample video group to obtain the fused sample video corresponding to the sample video group; wherein, the first video splicing operation is a time vertical splicing operation, and the video content of the fused sample video corresponding to the sample video group at any moment of its corresponding video duration includes the video content of each sample video in the sample video group at the corresponding moment. The pre-determined recognition model is trained based on all the fused sample videos to obtain the trained recognition model. The trained recognition model is used to identify whether there is abnormal behavior in the video content of any target video and, when abnormal behavior is identified in the video content of the target video, to determine at least one of the target abnormal behaviors in the video content of the target video.

2. The sample data processing method for model training according to claim 1, characterized in that, Before performing a first video stitching operation on all sample videos in each sample video group to obtain the fused sample video corresponding to that sample video group, the method further includes: For each sample video group, the duration of each sample video in the sample video group is calculated, and the sum of the durations of all sample videos in the sample video group is obtained to obtain the total video duration corresponding to the sample video group. It is then determined whether the total video duration corresponding to the sample video group is less than or equal to a predetermined video duration threshold. If the determination result is negative, the operation of performing a first video splicing operation on all sample videos in the sample video group to obtain the fused sample video corresponding to the sample video group is triggered. If the determination result is positive, a second video splicing operation is performed on all sample videos in the sample video group to obtain the fused sample video corresponding to the sample video group. The second video splicing operation is a time-level splicing operation, and the video content of the fused sample video corresponding to the sample video group at any moment of its corresponding video duration is the video content of one of the sample videos in the sample video group at the corresponding moment.

3. The sample data processing method for model training according to claim 1 or 2, characterized in that, After acquiring several sample videos to be trained, the method further includes: Determine whether there is at least one target sample video among all the sample videos, wherein the abnormal behavior annotation information corresponding to the target sample video includes two or more abnormal behavior categories; When it is determined that at least one of the target sample videos exists, the training abnormal behavior category corresponding to each target sample video is selected from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each target sample video according to the pre-determined abnormal behavior screening factor, and all abnormal behavior categories other than the selected training abnormal behavior categories are hidden from all abnormal behavior categories included in the abnormal behavior annotation information corresponding to each target sample video, so as to update the abnormal behavior annotation information corresponding to all the target sample videos.

4. The sample data processing method for model training according to claim 3, characterized in that, The sample videos are grouped according to the abnormal behavior categories of all target abnormal behaviors to be identified, resulting in several sample video groups, including: Based on the abnormal behavior categories of all target abnormal behaviors to be identified and the abnormal behavior annotation information corresponding to each sample video, all sample videos are divided into multiple sample video sets. Each sample video set includes multiple sample videos. One sample video set corresponds to one abnormal behavior category of the target abnormal behavior. The abnormal behavior annotation information corresponding to the sample videos in the same sample video set is the same, while the abnormal behavior annotation information corresponding to the sample videos in different sample video sets is different. One sample video is selected from each of the sample video sets, and the sample videos selected from each of the sample video sets are defined as a sample video group; Determine whether the number of all determined sample video groups reaches a first preset number threshold. If the determination result is no, repeat the operation of selecting one sample video from each sample video set and determining the sample video selected from each sample video set as a sample video group until the number of all determined sample video groups reaches the first preset number threshold.

5. The sample data processing method for model training according to claim 4, characterized in that, Before selecting one sample video from each of the sample video sets and determining the selected sample videos from each of the sample video sets as a sample video group, the method further includes: For any of the sample video sets, determine whether the total number of sample videos included in the sample video set is less than a second preset number threshold; when it is determined that the total number of sample videos included in the sample video set is less than the second preset number threshold, perform a sample video supplementation operation of the same abnormal behavior category on the sample video set according to the abnormal behavior category corresponding to the sample video set until the total number of sample videos included in the sample video set is not less than the second preset number threshold.

6. The sample data processing method for model training according to claim 4 or 5, characterized in that, The step of selecting one sample video from each of the sample video sets includes: Select one sample video from all unselected sample videos in one of the sample video sets as a benchmark sample video, obtain the video information corresponding to the benchmark sample video as benchmark information, and then filter sample videos from all unselected sample videos in the remaining sample video sets whose corresponding video information matches the benchmark information with a degree greater than or equal to a preset matching degree threshold.

7. A sample data processing device for model training, characterized in that, The apparatus is used to implement the sample data processing method for model training as described in any one of claims 1-6, and the apparatus comprises: The sample acquisition module is used to acquire several sample videos to be trained. The sample grouping module is used to group several sample videos according to the abnormal behavior category of all target abnormal behaviors to be identified, to obtain several sample video groups. Each sample video group includes at least two sample videos, and all abnormal behaviors involved in all sample videos included in each sample video group include all target abnormal behaviors. The sample fusion module is used to perform a video fusion operation on each of the sample video groups to obtain a fused sample video corresponding to each sample video group. The model training module is used to train a pre-determined recognition model based on all the fused sample videos to obtain a trained recognition model. The trained recognition model is used to identify whether there is abnormal behavior in the video content of any target video and, when abnormal behavior is identified in the video content of the target video, to determine at least one of the target abnormal behaviors in the video content of the target video.

8. A sample data processing device for model training, characterized in that, The device includes: Memory containing executable program code; A processor coupled to the memory; The processor calls the executable program code stored in the memory to execute the sample data processing method for model training as described in any one of claims 1-6.

9. A computer storage medium, characterized in that, The computer storage medium stores computer instructions, which, when invoked, are used to execute the sample data processing method for model training as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Image processing model training method based on artificial intelligence, and image processing method

    CN113821657A