Behavior recognition method, behavior recognition device, video detection method and storage medium
By performing timing analysis of the support sets and query sets in the behavior recognition task and determining the target category, the problem of high cost of large-sample behavior recognition and lack of scalability in the existing technology is solved, and the effect of improving the learning ability and scalability of small-sample behavior recognition methods is achieved.
Patent Information
- Application Number
- CN202210396744.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-04-15
AI Technical Summary
In the prior art, behavior recognition methods that rely on large samples lead to high recognition costs, while behavior recognition methods that rely on small samples lack scalability and are difficult to automatically identify new categories of behaviors.
By obtaining the support set and query set in the task to be identified, performing timing analysis, determining the target category to which the query set belongs from multiple candidate categories, improving the learning ability and scalability of the small sample behavior recognition method.
It is realized by performing timing analysis of the support set samples and query set samples in the recognition task, and determine the corresponding behavior category of the task to be recognized, thereby improving the learning ability and scalability of the small sample behavior recognition method and reducing the recognition cost.
Smart Images

Figure CN114817635B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a behavior recognition method, a behavior recognition device, a video detection method and a storage medium. Background Art
[0002] In the field of artificial intelligence, behavior recognition models trained with a large number of labeled samples are often used for behavior recognition. When a new category of behavior to be recognized appears, it is necessary to re-collect a large number of labeled samples based on the new category and re-train the behavior recognition model, which is costly. In response to this, technicians in related fields have been constantly trying various small sample behavior recognition methods in order to identify the new behavior category using a small number of labeled samples.
[0003] Compared with the behavior recognition method based on a large number of labeled samples, the small sample behavior recognition method provided by the related art has achieved significant improvement in behavior recognition performance. However, the small sample behavior recognition method provided by the related art has the following defects: poor learning ability, unable to be generalized to unseen domains, and thus difficult to automatically recognize newly added categories of behavior.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] The embodiments of the present invention provide a behavior recognition method, a behavior recognition device, a video detection method and a storage medium, so as to at least solve the technical problems in the related art that the behavior recognition method relying on large samples is prone to high recognition costs, while the behavior recognition method relying on small samples lacks scalability.
[0006] According to one aspect of an embodiment of the present invention, a behavior recognition method is provided, comprising: obtaining a task to be recognized, wherein the task to be recognized comprises: a support set and a query set, the support set being used to provide multiple candidate categories for behavior recognition of a video to be queried, and the query set being used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining a target category to which the query set belongs from multiple candidate categories.
[0007] According to another aspect of an embodiment of the present invention, a behavior recognition method is also provided, including: receiving a task to be recognized from a client, wherein the task to be recognized includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of a video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining a target category to which the query set belongs from multiple candidate categories; and feeding back the target category to the client.
[0008] According to another aspect of an embodiment of the present invention, a behavior recognition device is also provided, including: an acquisition module, used to acquire a task to be recognized, wherein the task to be recognized includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of a video to be queried, and the query set is used to determine the video to be queried; an identification module, used to perform time series analysis on the support set and the query set, and determine the target category to which the query set belongs from multiple candidate categories.
[0009] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, the computer-readable storage medium including a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned behavior recognition methods.
[0010] According to another aspect of an embodiment of the present invention, there is also provided an electronic device, comprising: a processor; and a memory, connected to the processor, for providing the processor with instructions for processing the following processing steps: obtaining a task to be identified, wherein the task to be identified comprises: a support set and a query set, the support set being used to provide multiple candidate categories for behavior recognition of a video to be queried, and the query set being used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining a target category to which the query set belongs from multiple candidate categories.
[0011] According to another aspect of an embodiment of the present invention, a video detection method is also provided, characterized in that it includes: obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories.
[0012] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of the small sample behavior recognition method, and further solving the technical problem in the related technology that the behavior recognition method relying on large samples is prone to cause high recognition cost, while the behavior recognition method relying on small samples lacks scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0014] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a behavior recognition method is shown;
[0015] Figure 2 is a flow chart of a behavior recognition method according to an embodiment of the present invention;
[0016] Figure 3 is a schematic diagram of an optional small sample behavior recognition process according to an embodiment of the present invention;
[0017] Figure 4 is a flow chart of another behavior recognition method according to an embodiment of the present invention;
[0018] Figure 5 is a schematic diagram of performing behavior recognition on a cloud server according to an embodiment of the present invention;
[0019] Figure 6 is a flow chart of a video detection method according to an embodiment of the present invention;
[0020] Figure 7 is a structural schematic diagram of a behavior recognition device according to an embodiment of the present invention;
[0021] Figure 8 is a schematic diagram of the structure of an optional behavior recognition device according to an embodiment of the present invention;
[0022] Fig. 9 is a structural block diagram of another computer terminal according to an embodiment of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] First, some nouns or terms that appear in the process of describing the embodiments of the present invention are subject to the following explanations:
[0026] Behavior recognition: refers to the process of analyzing the behavior category of the target object in the image.
[0027] Small-sample behavior recognition refers to the process of accurately identifying new categories of behaviors with only a few labeled samples. In small-sample behavior recognition, it is usually assumed that the base training set and the unlabeled target set are from the same domain.
[0028] Cross-domain small sample behavior recognition: refers to the process of more accurate small sample behavior recognition by reducing the domain shift when there is a domain shift between the basic training set and the unlabeled target set.
[0029] Self-attention: refers to the mechanism used to calculate the relationship between each unit in one sequence and all units in the other sequence when two identical sequences are input, and aggregate the relationship to improve the feature learning effect. For example: for a sentence, each word in the sentence should be related to other words in the sentence.
[0030] Episodic testing: refers to the process of testing in small sample behavior recognition tasks based on multiple episodes split by test samples, where the multiple episodes correspond to multiple tasks, and each task contains a support set (a set of labeled category samples) and a query set (a set of samples to be tested). Episodic testing can improve the performance of the model in small sample behavior recognition scenarios.
[0031] N-way-K-shot task: refers to a small-sample behavior recognition task in which the support set contains N behavior categories, and each behavior category in the N behavior categories corresponds to K samples (that is, the support set contains N×K samples).
[0032] Example 1
[0033] According to an embodiment of the present invention, a behavior recognition method embodiment is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0034] The method embodiment provided in the first embodiment of the present invention may be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG. 1 shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing a behavior recognition method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more (102a, 102b, ..., 102n are used to illustrate) processors 102 (the processor 102 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, a keyboard, a cursor control device (such as a mouse), an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. It can be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components as shown, or with Figure 1 Different configurations are shown.
[0035] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuits". The data processing circuits may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. In addition, the data processing circuit may be a single independent processing module, or may be incorporated in whole or in part into any of the other components in the computer terminal 10 (or mobile device). As involved in the embodiments of the present invention, the data processing circuit acts as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).
[0036] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the behavior recognition method in the embodiment of the present invention. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, realizing the above-mentioned behavior recognition method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely arranged relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0037] The transmission device 106 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one example, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0038] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0039] It should be noted that, in some optional embodiments, the above Figure 1 The computer device (or mobile device) shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. It should be noted that Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the above-described computer device (or mobile device).
[0040] Under the above operating environment, the present invention provides Figure 2 A behavior recognition method is shown. Figure 2 is a flow chart of a behavior recognition method according to an embodiment of the present invention. Figure 2 As shown, the behavior recognition method includes:
[0041] Step S21, obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of the video to be queried, and the query set is used to determine the video to be queried;
[0042] Step S22: performing a time series analysis on the support set and the query set, and determining the target category to which the query set belongs from a plurality of candidate categories.
[0043] In an embodiment of the present invention, the task to be identified may be a small sample behavior identification task. The task to be identified may be a task corresponding to one of multiple episodes in episodic testing. The task to be identified may include the support set and the query set.
[0044] The support set may be a set of labeled class samples, which may be used to provide multiple candidate classes for behavior recognition of a video to be queried. The display content of the video to be queried may include the behavior of a target object, and the multiple candidate classes may be multiple behavior classes to which the behavior of the target object may belong.
[0045] The query set may be a set of samples to be tested, and the query set may be used to determine the video to be queried. For example, the multiple samples to be tested included in the query set may be multiple video frames included in the video to be queried.
[0046] By performing a time series analysis on the support set and the query set in the above-mentioned identification task, the target category to which the query set belongs can be determined from the multiple candidate categories corresponding to the support set. The target category can be the behavior category to which the behavior action corresponding to each of the multiple test samples in the query set belongs.
[0047] Specifically, a time series analysis is performed on the support set and the query set to determine the target category to which the query set belongs from multiple candidate categories. Other method steps may also be included, and reference may be made to the further introduction to the embodiments of the present invention below, which will not be described in detail here.
[0048] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of the small sample behavior recognition method, and further solving the technical problem in the related technology that the behavior recognition method relying on large samples is prone to cause high recognition cost, while the behavior recognition method relying on small samples lacks scalability.
[0049] The above method of the embodiment of the present invention is further introduced below.
[0050] In an optional embodiment, in step S22, performing temporal analysis on the support set and the query set to determine the target category from multiple candidate categories includes the following method steps:
[0051] Step S221, using the target classification model to perform time series modeling on the support set to obtain a first time series dynamic relationship;
[0052] Step S222, determining a plurality of candidate categories based on the first temporal dynamic relationship;
[0053] Step S223, using the target classification model to perform time series modeling on the query set to obtain a second time series dynamic relationship;
[0054] Step S224, determining a target category from a plurality of candidate categories based on the first temporal dynamic relationship and the second temporal dynamic relationship;
[0055] The target classification model is trained using a sample data set, and the sample data set includes: a basic training set and an unlabeled target set, and the basic training set and the unlabeled target set come from different data domains.
[0056] In the above optional embodiment, a target classification model is used to perform time series analysis on the support set and the query set. The target classification model can be a neural network model trained by machine learning using a sample data set. The sample data set can include a basic training set and an unlabeled target set from different data domains.
[0057] By using the target classification model to perform temporal modeling on the support set of the task to be identified, the first temporal dynamic relationship can be obtained. The first temporal dynamic relationship can include the temporal relationship and dynamic relationship between multiple labeled category samples included in the support set.
[0058] For example, the temporal relationship between two labeled category samples may be the order relationship, time difference information, etc. of the two labeled category samples; the dynamic relationship between two labeled category samples may be the dynamic change categories, dynamic change paths, etc. between the behavior actions of the target objects corresponding to the two labeled category samples.
[0059] Based on the first temporal dynamic relationship, the multiple candidate categories may be determined, which may be multiple behavior categories to which the behavior of the target object displayed in the video to be queried may belong.
[0060] By using the target classification model to perform temporal modeling on the query set of the task to be identified, the second temporal dynamic relationship can be obtained. The second temporal dynamic relationship can include the temporal relationship and dynamic relationship between the multiple samples to be tested contained in the query set.
[0061] For example: the temporal relationship between two samples to be tested may be the sequence relationship, time difference information, etc. of the two samples to be tested; the dynamic relationship between two samples to be tested may be the dynamic change category, dynamic change path, etc. between the behavior actions of the target objects corresponding to the two samples to be tested.
[0062] Based on the first temporal dynamic relationship and the second temporal dynamic relationship, the target category is determined from the plurality of candidate categories. The target category may be a behavior category to which the behavior action corresponding to each of the plurality of samples to be tested in the query set belongs.
[0063] For example, the first temporal dynamic relationship and the corresponding second temporal dynamic relationship can be compared to obtain the difference between the first temporal dynamic relationship and the second temporal dynamic relationship, and then the target category to which the query set belongs can be determined from multiple candidate categories corresponding to the support set by analyzing the difference.
[0064] In an optional embodiment, in step S224, based on the first temporal dynamic relationship and the second temporal dynamic relationship, determining the target category from a plurality of candidate categories includes the following method steps:
[0065] Step S2241, based on the first time series dynamic relationship, performing mean calculation on each of the multiple candidate categories to obtain multiple calculation results, wherein the multiple calculation results are used to represent the sample means corresponding to the multiple candidate categories;
[0066] Step S2242, based on the multiple calculation results, performing difference analysis on the second time series dynamic relationship to obtain multiple analysis results, wherein the multiple analysis results are used to represent difference information between the multiple calculation results and the query set;
[0067] Step S2243, determining a target category from multiple candidate categories based on multiple analysis results.
[0068] In the above optional embodiment, the above first temporal dynamic relationship may include the temporal relationship and dynamic relationship between multiple labeled category samples included in the support set. The above second temporal dynamic relationship may include the temporal relationship and dynamic relationship between multiple samples to be tested included in the query set. The above multiple candidate categories may be multiple behavior categories to which the behavior actions of the target object displayed in the video to be queried may belong.
[0069] Based on the first temporal dynamic relationship, the mean calculation is performed on each of the multiple candidate categories. The sample mean of at least one labeled category sample corresponding to each of the multiple candidate categories is calculated according to the temporal relationship and dynamic relationship between the multiple labeled category samples included in the above support set.
[0070] Based on multiple calculation results, a difference analysis is performed on the second time series dynamic relationship. The method can be to calculate the distance between the sample mean of at least one labeled category sample corresponding to each candidate category in the multiple candidate categories and each sample to be tested in the multiple samples to be tested contained in the above query set under the second time series dynamic relationship, and perform a difference analysis based on the calculated distance to obtain multiple analysis results.
[0071] According to the multiple analysis results, a target category can be determined from the multiple candidate categories. The target category can be a candidate category corresponding to the minimum value of the difference data in the multiple analysis results. In an optional embodiment, the behavior recognition method further includes the following method steps:
[0072] Step S23, using the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and determining the first loss and the second loss respectively;
[0073] Step S24, jointly optimize the initial classification model based on the first loss and the second loss to obtain a target classification model.
[0074] In the above optional embodiment, the initial classification model can be an initial neural network model used to train a target classification model according to a sample data set. The sample data set includes a basic training set and an unlabeled target set. The basic training set and the unlabeled target set come from different data domains.
[0075] The first loss can be determined by using the initial classification model to perform time series modeling on the basic training set; the second loss can be determined by using the initial classification model to perform time series modeling on the unlabeled target set. The first loss and the second loss are used to optimize the neural network model.
[0076] The target classification model can be obtained by jointly optimizing the initial classification model based on the first loss and the second loss. The target classification model can be used to perform time series analysis on the support set and the query set, and determine the target category to which the query set belongs from multiple candidate categories corresponding to the support set.
[0077] For example, based on the initial classification model, the basic training set, and the unlabeled target set, the first loss Loss can be obtained. ce and the second loss ss Based on the first loss Loss ce and the second loss ss , an end-to-end training strategy is used to train the target classification model, and the neural network model is optimized based on the total loss Loss until the model converges. The calculation method of the total loss Loss in training can be shown in the following formula (1):
[0078] Loss=Loss ce +αLoss ss Formula (1)
[0079] In the above formula (1), α represents a balance coefficient.
[0080] In an optional embodiment, in step S23, the initial classification model is used to perform time series modeling on the basic training set and the unlabeled target set, and the first loss and the second loss are determined respectively, including the following method steps:
[0081] Step S231, using the initial classification model to perform time series modeling on the basic training set to obtain a third time series dynamic relationship;
[0082] Step S232, performing classification training based on the third temporal dynamic relationship to determine a first loss;
[0083] Step S233, using the initial classification model to perform time series modeling on the unlabeled target set to obtain a fourth time series dynamic relationship;
[0084] Step S234, performing comparative learning based on the third timing dynamic relationship and the fourth timing dynamic relationship to determine the second loss.
[0085] In the above optional embodiment, the initial classification model can be an initial neural network model used to train a target classification model according to a sample data set. The sample data set includes a basic training set and an unlabeled target set. The basic training set and the unlabeled target set come from different data domains.
[0086] By using the initial classification model to perform time series modeling on the basic training set, the third time series dynamic relationship can be obtained. The third time series dynamic relationship can include a time series relationship and a dynamic relationship between a plurality of labeled basic samples included in the basic training set. Classification training can be performed based on the third time series dynamic relationship to determine the first loss. The first loss can be a classification training loss.
[0087] By using the above-mentioned initial classification model to perform temporal modeling on the above-mentioned unlabeled target set, the above-mentioned fourth temporal dynamic relationship can be obtained. The fourth temporal dynamic relationship may include a temporal relationship and a dynamic relationship between a plurality of unlabeled target samples included in the unlabeled target set. Based on the above-mentioned third temporal dynamic relationship and the fourth temporal dynamic relationship, comparative learning is performed to determine the above-mentioned second loss. The second loss may be a comparative learning loss.
[0088] In an optional embodiment, in step S231, the initial classification model includes: a feature extraction stage, a time series modeling stage, and the initial classification model is used to perform time series modeling on the basic training set to obtain a third time series dynamic relationship, including the following method steps:
[0089] Step S2311, extracting features from the basic training set in the feature extraction phase to obtain a first feature sequence;
[0090] Step S2312, performing time series modeling on the first feature sequence in the time series modeling stage to obtain a third time series dynamic relationship.
[0091] In the above optional embodiment, the initial classification model can be an initial neural network model used to train the target classification model according to the sample data set. The sample data set includes a basic training set, which can be a set of multiple labeled basic samples. The initial classification model can include: a feature extraction stage and a time series modeling stage.
[0092] Specifically, in the feature extraction stage of the initial classification model, feature extraction can be performed on the basic training set to obtain a first feature sequence corresponding to the basic training set. The first feature sequence may include features corresponding to a plurality of labeled basic samples included in the basic training set.
[0093] Specifically, in the temporal modeling stage of the initial classification model, the first feature sequence can be temporally modeled to obtain the third temporal dynamic relationship. The third temporal dynamic relationship can include the temporal relationship and dynamic relationship between multiple labeled basic samples included in the basic training set.
[0094] Figure 3 FIG. 1 is a schematic diagram of an optional small sample behavior recognition process according to an embodiment of the present invention. Figure 3As shown, the sample data set used to train the small sample behavior recognition model (equivalent to the above-mentioned target classification model) includes labeled basic samples (equivalent to the above-mentioned basic training set).
[0095] For example, based on the sampled video Video0 with T video frames, a labeled video basic sample set can be determined: like Figure 3 As shown, the labeled video basic sample set V b Input the feature extraction network to extract features and get the feature sequence Among them, D represents the channel dimension of the feature sequence.
[0096] Still taking the training of a small sample behavior recognition model based on the sampled video Video0 as an example, Figure 3 As shown, the above feature sequence F b (equivalent to the first feature sequence mentioned above) is input into the temporal modeling sub-stage 1 in the temporal modeling stage to perform temporal modeling and obtain a labeled video basic sample set V b The correlation between multiple samples in is recorded as the time series dynamic relationship TM03 (equivalent to the third time series dynamic relationship mentioned above).
[0097] In an optional embodiment, in step S232, the initial classification model further includes: a supervised classification stage, performing classification training based on the third temporal dynamic relationship, and determining the first loss, including the following method steps:
[0098] Step S2321, performing supervised classification training on the third temporal dynamic relationship in the supervised classification stage to determine the first loss.
[0099] In the above optional embodiment, the initial classification model may be an initial neural network model used to train the target classification model according to the sample data set. The initial classification model may include: a supervised classification stage.
[0100] The third temporal dynamic relationship mentioned above may include a temporal relationship and a dynamic relationship between a plurality of labeled basic samples included in the basic training set.
[0101] The first loss may be a classification training loss. In the supervised classification stage in the initial classification model, supervised classification training is performed on the third temporal dynamic relationship to determine the classification training loss.
[0102] Still taking the training of a small sample behavior recognition model based on the sampled video Video0 as an example, Figure 3As shown in FIG. 1 , when training a small sample behavior recognition model, in the supervised classification stage of the initial neural network model, the temporal dynamic relationship TM03 (equivalent to the third temporal dynamic relationship mentioned above) output by the temporal modeling sub-stage 1 is received; based on the temporal dynamic relationship TM03, multiple multi-layer perceptrons (MLPs) are used to classify the labeled video basic sample set V b Classify and then determine the classification training loss Loss ce (equivalent to the first loss mentioned above).
[0103] Specifically, the method for supervised classification training based on the temporal dynamic relationship TM03 can be shown as follows:
[0104]
[0105]
[0106]
[0107] In the above formulas (2) to (4), y represents the actual behavior category, Represents the predicted category output by the supervised classification stage.
[0108] In an optional embodiment, in step S233, the initial classification model is used to perform time series modeling on the unlabeled target set to obtain a fourth time series dynamic relationship, including the following method steps:
[0109] Step S2331, performing data augmentation processing on the unlabeled target set to obtain a processing result;
[0110] Step S2332, performing feature extraction on the processing result in the feature extraction stage to obtain a second feature sequence;
[0111] Step S2333, performing time series modeling on the second feature sequence in the time series modeling stage to obtain a fourth time series dynamic relationship.
[0112] In the above optional embodiment, the initial classification model may be an initial neural network model for training a target classification model according to a sample data set. The sample data set includes an unlabeled target set, which may be a collection of multiple unlabeled target samples. The initial classification model may include: a feature extraction stage and a time series modeling stage.
[0113] Performing data augmentation processing on the unlabeled target sets to obtain processing results may be: performing data augmentation processing on the unlabeled target sets at least once respectively to obtain at least one processing result.
[0114] Specifically, in the feature extraction stage of the initial classification model, feature extraction can be performed on the above processing results to obtain a second feature sequence corresponding to the unlabeled target set. The second feature sequence can include features corresponding to multiple unlabeled target samples included in the unlabeled target set.
[0115] Specifically, in the feature extraction stage of the initial classification model, feature extraction can be performed on the at least one processing result to obtain at least one second feature sequence corresponding to the unlabeled target set. Each of the at least one second feature sequence can include features corresponding to multiple unlabeled target samples included in the unlabeled target set.
[0116] Specifically, in the temporal modeling stage of the initial classification model, the second feature sequence can be temporally modeled to obtain the fourth temporal dynamic relationship. The fourth temporal dynamic relationship can include the temporal relationship and dynamic relationship between multiple unlabeled target samples included in the unlabeled target set.
[0117] Still taking the training of a small sample behavior recognition model based on the sampled video Video0 as an example, Figure 3 As shown in FIG. 1 , the sample data set used to train the small sample behavior recognition model (equivalent to the above target classification model) also includes unlabeled target samples (equivalent to the above unlabeled target set). There is a significant domain difference between the above labeled basic samples and the unlabeled target samples. Based on the sampled video Video0 with T video frames, an unlabeled video target sample set can be determined:
[0118] Still like Figure 3 As shown, the unlabeled video target sample set V u After two data augmentations (such as Figure 3 After the data augmentation 1 and data augmentation 2 shown in the figure are input into the feature extraction network for feature extraction, the corresponding two feature sequences can be obtained: and Among them, D represents the channel dimension of the feature sequence.
[0119] Still like Figure 3 As shown, the above feature sequence F u1 and F u2 (equivalent to the second feature sequence) are respectively input into the temporal modeling sub-stage 2 and the temporal modeling sub-stage 3 in the temporal modeling stage, and temporal modeling is performed respectively to obtain an unlabeled video target sample set V u The correlation between multiple samples in is recorded as the time series dynamic relationship TM04 (equivalent to the fourth time series dynamic relationship mentioned above).
[0120] It is easy to notice that in the time series modeling stage, the multi-head self-attention mechanism (MSA) is used for time series modeling. Specifically, the feature sequence F b Input time series modeling sub-stage 1, the feature sequence F u1 Input time series modeling sub-stage 2, the feature sequence F u2 Enter the timing modeling sub-stage 3; perform timing modeling in the three timing modeling sub-stages respectively.
[0121] Specifically, the above method of using the multi-head self-attention mechanism for time series modeling can be shown as follows:
[0122] X′=Norm(F) Formula (5)
[0123]
[0124] H i =Attention(Q i ; K i ; V i ) Formula (7)
[0125]
[0126] In the above formulas (5) to (8), Q i , K i and V i represents the fully connected layer operation of the neural network model, Q i =X′W i q , K i =X′W i k , V i =X′W i v , where W i q , W i q and W i q is the fully connected layer parameter of the neural network model, and N head Indicates the number of multi-head attention (usually set to N head =8); Indicates the scaling parameter (usually set d k is the feature channel dimension).
[0127] It is easy to notice that by using unlabeled target samples to train the small sample behavior recognition model, the domain differences between sample data sets can be reduced, thereby making the trained small sample behavior recognition model more scalable.
[0128] It is easy to notice that by using the multi-head self-attention mechanism for time series modeling in the time series modeling stage, the distance between each two features in multiple features is reduced, which facilitates the modeling of non-local relationships between features. This allows the model to learn different information in subspaces of different representations, increases the number of relationships that the model can pay attention to, and thus improves the robustness and expressiveness of the model; it can also improve the model's ability to learn long-range time series dependencies, thereby enabling the model to more fully explore the correlation between multiple video samples.
[0129] It is easy to notice that when using the multi-head self-attention mechanism for time series modeling, by using scaling parameters, the value of the obtained feature can be controlled to be within an appropriate range (not too large), thereby improving the stability and convergence speed of model training.
[0130] In an optional embodiment, in step S234, the initial classification model further includes: a self-supervised contrastive learning stage, performing contrastive learning based on the third temporal dynamic relationship and the fourth temporal dynamic relationship to determine the second loss, including the following method steps:
[0131] Step S2341, performing self-supervised comparative learning on the third temporal dynamic relationship and the fourth temporal dynamic relationship in the self-supervised comparative learning stage to determine the second loss.
[0132] In the above optional embodiment, the initial classification model may be an initial neural network model used to train the target classification model according to the sample data set. The initial classification model may include: a self-supervised contrast learning stage.
[0133] The third temporal dynamic relationship may include the temporal relationship and dynamic relationship between multiple labeled basic samples included in the basic training set. The fourth temporal dynamic relationship may include the temporal relationship and dynamic relationship between multiple unlabeled target samples included in the unlabeled target set.
[0134] The first loss may be a contrastive learning loss. In the self-supervised contrastive learning stage in the initial classification model, the third temporal dynamic relationship and the fourth temporal dynamic relationship are subjected to self-supervised contrastive learning to determine the contrastive learning loss.
[0135] Still taking the training of a small sample behavior recognition model based on the sampled video Video0 as an example, Figure 3As shown, when training the small sample behavior recognition model, in the self-supervised contrast learning stage of the initial neural network model, the temporal dynamic relationship TM03 and the temporal dynamic relationship TM04 (equivalent to the fourth temporal dynamic relationship mentioned above) outputted in the temporal modeling stage are received; based on the temporal dynamic relationship TM03 and the temporal dynamic relationship TM04, the labeled video basic sample set V is trained. b and unlabeled video target sample set V u Perform contrastive learning to determine the contrastive learning loss Loss ss (equivalent to the second loss mentioned above).
[0136] The contrastive learning method is one of the intuitive learning methods in self-supervised learning. It is similar to the way humans can identify cats even when the concept of cats is unknown (even without a linguistic definition of cats) by comparing the same features between different individuals of the same cat species and the different features between cats and different individuals of different species.
[0137] In the process of contrastive learning, the discriminability of instances is enhanced by pulling positive sample pairs closer and pushing negative sample pairs further away in the feature space.
[0138] Still taking the training of a small sample behavior recognition model based on the sampled video Video0 as an example, Figure 3 As shown, in the self-supervised contrastive learning stage of the initial neural network model, the labeled video base sample set V b and unlabeled video target sample set V u The method for contrastive learning can be: calculate the labeled video base sample set V b and unlabeled video target sample set V u The similarity of the positive sample pairs in the positive sample pair is increased; the labeled video basic sample set V is calculated. b and unlabeled video target sample set V u The similarity of the negative sample pairs in the above example is reduced by the similarity of the negative sample pairs. Specifically, the above contrastive learning method can be shown as the following formula (9) to formula (10):
[0139]
[0140]
[0141] In the above formulas (9) to (10), τ represents the temperature coefficient, N represents the number of sampled videos in a training batch, Representation characteristics and Features The similarity between (usually using cosine similarity), where the feature It can be a labeled video base sample set Vb The corresponding feature sequence F b Any of the features, features It can be an unlabeled video target sample set V u The corresponding feature sequence F u1 and F u2 any one of the features.
[0142] It is easy to notice that through self-supervised contrastive learning, the generalization of the representation learned by the model can be improved, so that the trained small sample behavior recognition model can have more effective domain transfer capabilities and improve the recognition performance of the small sample behavior recognition model in the target domain.
[0143] It is easy to notice that the method provided by the embodiment of the present invention can improve the efficiency and accuracy of cross-domain behavior recognition in small sample behavior recognition scenarios; it can also use the time series modeling process to extract the time series dynamic relationship in the video to be queried, and combine self-supervised comparative learning to improve the generalization performance of the model, thereby reducing the difficulty of migration from the training data domain to the target test domain during the behavior recognition process.
[0144] It should be noted that the focus of the method provided by the present invention is: designing a temporal modeling process based on a self-attention mechanism to more fully explore the temporal dynamic relationship between samples, thereby improving the effectiveness of sample expression content; and through a self-supervised comparative learning process, more fully explore the representation of unlabeled data samples to enhance the migration ability in the model behavior recognition process, and achieve effective generalization in cross-domain scenarios. However, the present invention does not limit other non-key implementation methods.
[0145] One embodiment of the present invention further provides a behavior recognition method, which is run on a cloud server. Figure 4 is a flow chart of another behavior recognition method according to an embodiment of the present invention. Figure 4 As shown, the behavior recognition method includes:
[0146] Step S41, receiving a task to be identified from a client, wherein the task to be identified includes: a support set and a query set, wherein the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried;
[0147] Step S42, performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories;
[0148] Step S43: Feedback the target category to the client.
[0149] Optionally, Figure 5 is a schematic diagram of performing behavior recognition on a cloud server according to an embodiment of the present invention, such as Figure 5 As shown in the figure, the client uploads the task to be identified to the cloud server, where the task to be identified includes: a support set and a query set. The support set is used to provide multiple candidate categories for the behavior identification of the video to be queried, and the query set is used to determine the video to be queried; the cloud server performs time series analysis on the support set and the query set, and determines the target category to which the query set belongs from multiple candidate categories. Then, the cloud server will feedback the target category to the above client, and the final target category will be provided to the user through the client's graphical user interface.
[0150] It should be noted that the above-mentioned behavior recognition method provided in the embodiment of the present invention can be applicable to, but is not limited to, any practical application scenario involving behavior recognition. By means of interaction between the SaaS server and the client, the target category corresponding to the task to be recognized is determined by performing time series analysis on the support set samples and the query set samples in the task to be recognized, and the returned target category is provided to the user through the client.
[0151] One embodiment of the present invention further provides a video detection method. Figure 6 is a flow chart according to an embodiment of the present invention, such as Figure 6 As shown, the video detection method includes:
[0152] Step S61, obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of the video to be queried, and the query set is used to determine the video to be queried;
[0153] Step S62: Perform time series analysis on the support set and the query set to determine the target category to which the query set belongs from multiple candidate categories.
[0154] In the above-mentioned video detection method, the above-mentioned task to be identified may be a small sample behavior recognition task based on the video to be queried. The task to be identified based on the video to be queried may be a task corresponding to one of the multiple episodes in Episodic testing. The task to be identified based on the video to be queried may include the above-mentioned support set and the above-mentioned query set.
[0155] The support set may be a set of labeled video samples, which may be used to provide multiple candidate categories for behavior recognition of a video to be queried. The display content of the video to be queried may include the behavior of a target object, and the multiple candidate categories may be multiple behavior categories to which the behavior of the target object may belong.
[0156] The query set may be a set of video samples to be tested, and the query set may be used to determine the video to be queried. For example, the multiple video samples to be tested included in the query set may be multiple video frames included in the video to be queried.
[0157] By performing a time series analysis on the support set and the query set in the above-mentioned identification task based on the query video, the target category to which the query set belongs can be determined from the multiple candidate categories corresponding to the support set. The target category can be the behavior category to which the behavior action corresponding to each of the multiple video samples to be tested in the query set belongs.
[0158] Specifically, a time series analysis is performed on the support set and the query set to determine the target category to which the query set belongs from multiple candidate categories. Other method steps may also be included, and reference may be made to the introduction of the corresponding content in the behavior recognition method above, which will not be repeated here.
[0159] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of a small sample behavior recognition method based on the video to be queried, and further solving the technical problem in the related technology that a video recognition method relying on a large sample is prone to cause a high video recognition cost, while a video recognition method relying on a small sample lacks scalability.
[0160] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0161] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention.
[0162] Example 2
[0163] According to an embodiment of the present invention, a device for implementing the above behavior recognition method is also provided. Figure 7 is a schematic diagram of a behavior recognition device according to an embodiment of the present invention. Figure 7 As shown, the device includes: an acquisition module 71 and an identification module 72, wherein:
[0164] The acquisition module 71 is used to acquire the task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of the video to be queried, and the query set is used to determine the video to be queried; the identification module 72 is used to perform time series analysis on the support set and the query set, and determine the target category to which the query set belongs from multiple candidate categories.
[0165] Optionally, the above-mentioned identification module 72 is also used to: use the target classification model to perform time series modeling on the support set to obtain a first time series dynamic relationship; determine multiple candidate categories based on the first time series dynamic relationship; use the target classification model to perform time series modeling on the query set to obtain a second time series dynamic relationship; based on the first time series dynamic relationship and the second time series dynamic relationship, determine the target category from multiple candidate categories; wherein the target classification model is trained using a sample data set, and the sample data set includes: a basic training set and an unlabeled target set, and the basic training set and the unlabeled target set come from different data domains.
[0166] Optionally, the above-mentioned identification module 72 is also used to: based on the first time series dynamic relationship, perform mean calculation on each candidate category in multiple candidate categories to obtain multiple calculation results, wherein the multiple calculation results are used to represent the sample means corresponding to the multiple candidate categories; based on the multiple calculation results, perform difference analysis on the second time series dynamic relationship to obtain multiple analysis results, wherein the multiple analysis results are used to represent the difference information between the multiple calculation results and the query set; and determine the target category from the multiple candidate categories based on the multiple analysis results.
[0167] Optionally, Figure 8 is a schematic diagram of the structure of an optional behavior recognition device according to an embodiment of the present invention. Figure 8 As shown, the device includes Figure 7 In addition to all the modules shown, it also includes: a modeling module 73, which is used to use the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and determine the first loss and the second loss respectively; an optimization module 74, which is used to jointly optimize the initial classification model based on the first loss and the second loss to obtain a target classification model.
[0168] Optionally, the above-mentioned modeling module 73 is also used to: use the initial classification model to perform time series modeling on the basic training set to obtain a third time series dynamic relationship; perform classification training based on the third time series dynamic relationship to determine the first loss; use the initial classification model to perform time series modeling on the unlabeled target set to obtain a fourth time series dynamic relationship; perform comparative learning based on the third time series dynamic relationship and the fourth time series dynamic relationship to determine the second loss.
[0169] Optionally, the initial classification model includes: a feature extraction stage and a time series modeling stage. The above-mentioned modeling module 73 is also used to: perform feature extraction on the basic training set in the feature extraction stage to obtain a first feature sequence; perform time series modeling on the first feature sequence in the time series modeling stage to obtain a third time series dynamic relationship.
[0170] Optionally, the initial classification model further includes: a supervised classification stage, and the above-mentioned modeling module 73 is also used to: perform supervised classification training on the third temporal dynamic relationship in the supervised classification stage to determine the first loss.
[0171] Optionally, the above-mentioned modeling module 73 is also used to: perform data augmentation processing on the unlabeled target set to obtain a processing result; perform feature extraction on the processing result in the feature extraction stage to obtain a second feature sequence; perform time series modeling on the second feature sequence in the time series modeling stage to obtain a fourth time series dynamic relationship.
[0172] Optionally, the initial classification model also includes: a self-supervised contrastive learning stage, and the above-mentioned modeling module 73 is also used to: perform self-supervised contrastive learning on the third temporal dynamic relationship and the fourth temporal dynamic relationship in the self-supervised contrastive learning stage to determine the second loss.
[0173] It should be noted that the acquisition module 71 and the identification module 72 correspond to steps S21 to S22 in Example 1, and the examples and application scenarios implemented by the two modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above-mentioned Example 1. It should be noted that the above-mentioned modules, as part of the device, can be run in the computer terminal 10 provided in Example 1.
[0174] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of the small sample behavior recognition method, and further solving the technical problem in the related technology that the behavior recognition method relying on large samples is prone to cause high recognition cost, while the behavior recognition method relying on small samples lacks scalability.
[0175] It should be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1, which will not be repeated here.
[0176] Example 3
[0177] According to an embodiment of the present invention, an embodiment of an electronic device is also provided, and the electronic device may be any computing device in a computing device group. The electronic device includes: a processor and a memory, wherein:
[0178] A memory is connected to the processor and is used to provide instructions for the processor to process the following processing steps: obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories.
[0179] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of the small sample behavior recognition method, and further solving the technical problem in the related technology that the behavior recognition method relying on large samples is prone to cause high recognition cost, while the behavior recognition method relying on small samples lacks scalability.
[0180] It should be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1, which will not be repeated here.
[0181] Example 4
[0182] The embodiment of the present invention can provide a computer terminal, which can be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the computer terminal can also be replaced by a terminal device such as a mobile terminal.
[0183] Optionally, in this embodiment, the computer terminal may be located in at least one network device among a plurality of network devices of the computer network.
[0184] In this embodiment, the above-mentioned computer terminal can execute the program code of the following steps in the behavior recognition method: obtaining the task to be recognized, wherein the task to be recognized includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories.
[0185] Optionally, Fig. 9 is a structural block diagram of another computer terminal according to an embodiment of the present invention. Fig. 9 As shown, the computer terminal may include: one or more (only one is shown in the figure) processors 122 , a memory 124 , and a peripheral interface 126 .
[0186] Among them, the memory can be used to store software programs and modules, such as program instructions / modules corresponding to the behavior recognition method and device in the embodiment of the present invention. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, realizing the above-mentioned behavior recognition method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include a memory remotely arranged relative to the processor, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0187] The processor can call the information and application programs stored in the memory through the transmission device to perform the following steps: obtaining the task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories.
[0188] Optionally, the processor may also execute program code for the following steps: performing time series modeling on the support set using the target classification model to obtain a first time series dynamic relationship; determining a plurality of candidate categories based on the first time series dynamic relationship; performing time series modeling on the query set using the target classification model to obtain a second time series dynamic relationship; determining a target category from a plurality of candidate categories based on the first time series dynamic relationship and the second time series dynamic relationship; wherein the target classification model is trained using a sample data set, and the sample data set includes: a basic training set and an unlabeled target set, and the basic training set and the unlabeled target set come from different data domains.
[0189] Optionally, the processor may also execute program code of the following steps: based on the first time series dynamic relationship, performing mean calculation on each of multiple candidate categories to obtain multiple calculation results, wherein the multiple calculation results are used to represent the sample means corresponding to the multiple candidate categories; based on the multiple calculation results, performing difference analysis on the second time series dynamic relationship to obtain multiple analysis results, wherein the multiple analysis results are used to represent the difference information between the multiple calculation results and the query set; and determining the target category from the multiple candidate categories based on the multiple analysis results.
[0190] Optionally, the processor may also execute the program code of the following steps: using the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and determining the first loss and the second loss respectively; and jointly optimizing the initial classification model based on the first loss and the second loss to obtain the target classification model.
[0191] Optionally, the processor may also execute the program code of the following steps: performing time series modeling on the basic training set using the initial classification model to obtain a third time series dynamic relationship; performing classification training based on the third time series dynamic relationship to determine a first loss; performing time series modeling on the unlabeled target set using the initial classification model to obtain a fourth time series dynamic relationship; performing comparative learning based on the third time series dynamic relationship and the fourth time series dynamic relationship to determine a second loss.
[0192] Optionally, the processor may also execute program codes of the following steps: extracting features from a basic training set in a feature extraction phase to obtain a first feature sequence; and performing time series modeling on the first feature sequence in a time series modeling phase to obtain a third time series dynamic relationship.
[0193] Optionally, the processor may also execute program code of the following steps: performing supervised classification training on the third temporal dynamic relationship in a supervised classification stage to determine a first loss.
[0194] Optionally, the processor may also execute the program code of the following steps: performing data augmentation processing on the unlabeled target set to obtain a processing result; performing feature extraction on the processing result in the feature extraction stage to obtain a second feature sequence; performing time series modeling on the second feature sequence in the time series modeling stage to obtain a fourth time series dynamic relationship.
[0195] Optionally, the processor may also execute program code of the following steps: performing self-supervised comparative learning on the third temporal dynamic relationship and the fourth temporal dynamic relationship in a self-supervised comparative learning stage to determine a second loss.
[0196] The processor can call the information and application stored in the memory through the transmission device to perform the following steps: receiving the task to be identified from the client, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for the behavior identification of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories; and feeding back the target category to the client.
[0197] In an embodiment of the present invention, a task to be identified is obtained, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried, and a target category to which the query set belongs is determined from multiple candidate categories by performing time series analysis on the support set and the query set, thereby achieving the purpose of determining the behavior category corresponding to the task to be identified by performing time series analysis on the support set samples and the query set samples in the task to be identified, thereby achieving the technical effect of improving the learning ability and scalability of the small sample behavior recognition method, and further solving the technical problem in the related technology that the behavior recognition method relying on large samples is prone to cause high recognition cost, while the behavior recognition method relying on small samples lacks scalability.
[0198] It can be understood by those skilled in the art that Fig. 9 The structure shown is for illustration only, and the computer terminal may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (Mobile Internet Devices, MID), a PAD, or other terminal devices. Fig. 9 The structure of the electronic device is not limited. For example, the computer terminal may also include Fig. 9 More or fewer components (such as network interfaces, display devices, etc.) shown in, or having Fig. 9 Different configurations are shown.
[0199] A person of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0200] According to an embodiment of the present invention, an embodiment of a storage medium is also provided. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the behavior recognition method provided in the above embodiment 1.
[0201] Optionally, in this embodiment, the above storage medium may be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group.
[0202] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining the target category to which the query set belongs from multiple candidate categories.
[0203] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: using the target classification model to perform time series modeling on the support set to obtain a first time series dynamic relationship; determining multiple candidate categories based on the first time series dynamic relationship; using the target classification model to perform time series modeling on the query set to obtain a second time series dynamic relationship; based on the first time series dynamic relationship and the second time series dynamic relationship, determining a target category from multiple candidate categories; wherein the target classification model is trained using a sample data set, and the sample data set includes: a basic training set and an unlabeled target set, and the basic training set and the unlabeled target set come from different data domains.
[0204] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: based on the first time series dynamic relationship, performing mean calculation on each of the multiple candidate categories to obtain multiple calculation results, wherein the multiple calculation results are used to represent the sample means corresponding to the multiple candidate categories; based on the multiple calculation results, performing difference analysis on the second time series dynamic relationship to obtain multiple analysis results, wherein the multiple analysis results are used to represent the difference information between the multiple calculation results and the query set; and determining the target category from the multiple candidate categories based on the multiple analysis results.
[0205] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: using the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and determining the first loss and the second loss respectively; based on the first loss and the second loss, jointly optimizing the initial classification model to obtain the target classification model.
[0206] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: performing time series modeling on a basic training set using an initial classification model to obtain a third time series dynamic relationship; performing classification training based on the third time series dynamic relationship to determine a first loss; performing time series modeling on an unlabeled target set using an initial classification model to obtain a fourth time series dynamic relationship; performing comparative learning based on the third time series dynamic relationship and the fourth time series dynamic relationship to determine a second loss.
[0207] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: performing feature extraction on a basic training set in a feature extraction phase to obtain a first feature sequence; performing time series modeling on the first feature sequence in a time series modeling phase to obtain a third time series dynamic relationship.
[0208] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: performing supervised classification training on the third temporal dynamic relationship in a supervised classification stage to determine a first loss.
[0209] Optionally, in this embodiment, the storage medium is configured to store program codes for executing the following steps: performing data augmentation processing on the unlabeled target set to obtain a processing result; performing feature extraction on the processing result in a feature extraction stage to obtain a second feature sequence; and performing time series modeling on the second feature sequence in a time series modeling stage to obtain a fourth time series dynamic relationship.
[0210] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: performing self-supervised contrastive learning on the third temporal dynamic relationship and the fourth temporal dynamic relationship in a self-supervised contrastive learning stage to determine the second loss.
[0211] Optionally, in this embodiment, the storage medium is configured to store program code for executing the following steps: receiving a task to be identified from a client, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior recognition of the video to be queried, and the query set is used to determine the video to be queried; performing time series analysis on the support set and the query set, and determining a target category to which the query set belongs from multiple candidate categories; and feeding back the target category to the client.
[0212] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0213] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0214] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0215] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0216] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0217] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0218] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A behavior recognition method, characterized in that: include: Acquire a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried; Performing temporal modeling on the support set to obtain a first temporal dynamic relationship, and performing temporal modeling on the query set to obtain a second temporal dynamic relationship, wherein the first temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of labeled category samples included in the support set, and the second temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of samples to be tested included in the query set; Based on the first temporal dynamic relationship and the second temporal dynamic relationship, a target category to which the query set belongs is determined from the multiple candidate categories, wherein the target category is determined based on a difference result between the first temporal dynamic relationship and the second temporal dynamic relationship.
2. The behavior recognition method according to claim 1, characterized in that: Performing time series modeling on the support set to obtain the first time series dynamic relationship, and performing time series modeling on the query set to obtain the second time series dynamic relationship include: Using a target classification model to perform time series modeling on the support set to obtain the first time series dynamic relationship; Determine the plurality of candidate categories based on the first temporal dynamic relationship; Using the target classification model to perform time series modeling on the query set to obtain the second time series dynamic relationship; The target classification model is trained using a sample data set, and the sample data set includes: a basic training set and an unlabeled target set, and the basic training set and the unlabeled target set are from different data domains.
3. The behavior recognition method according to claim 2, characterized in that: Based on the first temporal dynamic relationship and the second temporal dynamic relationship, determining the target category from the plurality of candidate categories comprises: Based on the first time series dynamic relationship, perform mean calculation on each of the multiple candidate categories to obtain multiple calculation results, wherein the multiple calculation results are used to represent sample means corresponding to the multiple candidate categories; Based on the multiple calculation results, performing difference analysis on the second time series dynamic relationship to obtain multiple analysis results, wherein the multiple analysis results are used to represent difference information between the multiple calculation results and the query set; The target category is determined from the plurality of candidate categories according to the plurality of analysis results.
4. The behavior recognition method according to claim 2, characterized in that: The method further comprises: Using the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and determining a first loss and a second loss respectively; The initial classification model is jointly optimized based on the first loss and the second loss to obtain the target classification model.
5. The behavior recognition method according to claim 4, characterized in that: Using the initial classification model to perform time series modeling on the basic training set and the unlabeled target set, and respectively determining the first loss and the second loss includes: Using the initial classification model to perform time series modeling on the basic training set to obtain a third time series dynamic relationship; Perform classification training based on the third temporal dynamic relationship to determine the first loss; Using the initial classification model to perform time series modeling on the unlabeled target set to obtain a fourth time series dynamic relationship; The second loss is determined by performing comparative learning based on the third timing dynamic relationship and the fourth timing dynamic relationship.
6. The behavior recognition method according to claim 5, characterized in that: The initial classification model includes: a feature extraction stage and a time series modeling stage. The initial classification model is used to perform time series modeling on the basic training set to obtain the third time series dynamic relationship, which includes: In the feature extraction stage, feature extraction is performed on the basic training set to obtain a first feature sequence; In the time series modeling stage, time series modeling is performed on the first feature sequence to obtain the third time series dynamic relationship.
7. The behavior recognition method according to claim 6, characterized in that: The initial classification model further includes: a supervised classification stage, performing classification training based on the third temporal dynamic relationship, and determining the first loss includes: In the supervised classification stage, supervised classification training is performed on the third temporal dynamic relationship to determine the first loss.
8. The behavior recognition method according to claim 6, characterized in that: The initial classification model is used to perform time series modeling on the unlabeled target set to obtain the fourth time series dynamic relationship, which includes: Performing data augmentation processing on the unlabeled target set to obtain a processing result; In the feature extraction stage, feature extraction is performed on the processing result to obtain a second feature sequence; In the timing modeling stage, timing modeling is performed on the second feature sequence to obtain the fourth timing dynamic relationship.
9. The behavior recognition method according to claim 8, characterized in that: The initial classification model further includes: a self-supervised contrastive learning stage, performing contrastive learning based on the third temporal dynamic relationship and the fourth temporal dynamic relationship, and determining the second loss includes: In the self-supervised contrastive learning stage, self-supervised contrastive learning is performed on the third temporal dynamic relationship and the fourth temporal dynamic relationship to determine the second loss.
10. A behavior recognition method, characterized in that: include: Receiving a task to be identified from a client, wherein the task to be identified includes: a support set and a query set, wherein the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried; Performing temporal modeling on the support set to obtain a first temporal dynamic relationship, and performing temporal modeling on the query set to obtain a second temporal dynamic relationship, wherein the first temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of labeled category samples included in the support set, and the second temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of samples to be tested included in the query set; Based on the first temporal dynamic relationship and the second temporal dynamic relationship, determining a target category to which the query set belongs from the multiple candidate categories, wherein the target category is determined based on a difference result between the first temporal dynamic relationship and the second temporal dynamic relationship; The target category is fed back to the client.
11. A behavior recognition device, characterized in that: include: An acquisition module, used for acquiring a task to be identified, wherein the task to be identified includes: a support set and a query set, wherein the support set is used for providing a plurality of candidate categories for behavior identification of a video to be queried, and the query set is used for determining the video to be queried; An identification module is used to: perform temporal modeling on the support set to obtain a first temporal dynamic relationship, and perform temporal modeling on the query set to obtain a second temporal dynamic relationship, wherein the first temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between multiple labeled category samples included in the support set, and the second temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between multiple samples to be tested included in the query set; based on the first temporal dynamic relationship and the second temporal dynamic relationship, determine the target category to which the query set belongs from the multiple candidate categories, wherein the target category is determined based on the difference between the first temporal dynamic relationship and the second temporal dynamic relationship.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the behavior recognition method described in any one of claims 1 to 10.
13. An electronic device, characterized in that: include: processor; as well as A memory, connected to the processor, configured to provide the processor with instructions for processing the following processing steps: Step 1, obtaining a task to be identified, wherein the task to be identified includes: a support set and a query set, wherein the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried; Step 2: Performing a time series analysis on the support set and the query set, and determining the target category to which the query set belongs from the multiple candidate categories.
14. A video detection method, characterized in that: include: Acquire a task to be identified, wherein the task to be identified includes: a support set and a query set, the support set is used to provide multiple candidate categories for behavior identification of a video to be queried, and the query set is used to determine the video to be queried; Performing temporal modeling on the support set to obtain a first temporal dynamic relationship, and performing temporal modeling on the query set to obtain a second temporal dynamic relationship, wherein the first temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of labeled category samples included in the support set, and the second temporal dynamic relationship is used to represent the temporal relationship and dynamic relationship between a plurality of samples to be tested included in the query set; Based on the first temporal dynamic relationship and the second temporal dynamic relationship, a target category to which the query set belongs is determined from the multiple candidate categories, wherein the target category is determined based on a difference result between the first temporal dynamic relationship and the second temporal dynamic relationship.
Citation Information
Patent Citations
Small sample action recognition model training method and device, electronic equipment and storage medium
CN114282047A