Video analysis method, device, electronic device and storage medium

By using frame-level annotation in the video analysis model, the model parameters are adjusted based on the initial probability and training loss of the sample video clip, and the problem of low positioning accuracy in the existing technology is solved, and the accuracy of video analysis is improved.

CN115115985BActive Publication Date: 2025-08-15TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210746923.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-28
Publication Date
2025-08-15
Estimated Expiration
2042-06-28

AI Technical Summary

Technical Problem

The existing video analysis model has fewer learning features and low positioning accuracy and accuracy due to the small number of sample labels marked by sample videos.

Method used

By determining the sample reference probability based on the initial probability of each video clip of the sample video, the sample reference probability is determined, and parameter adjustments are made to the video analysis model based on the first training loss and the second training loss, the video analysis model is realized to achieve frame-level annotation and improve the training effect of the model.

Benefits of technology

Without increasing manpower consumption, the positioning accuracy and accuracy of the video analysis model are significantly improved, and the overall accuracy of video analysis is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115985B_ABST
    Figure CN115115985B_ABST
Patent Text Reader

Abstract

The present application relates to the field of video analysis technology, and in particular to a video analysis method, device, electronic device and storage medium for improving the accuracy of video analysis. The method comprises: based on a trained target video analysis model, obtaining the initial probability that each video clip contains each set event, based on each initial probability and a preset probability threshold, obtaining the positioning information of the target event contained in the video to be analyzed, and based on the positioning information, analyzing the target event contained in the video to be analyzed; wherein, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss, the first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is obtained based on the reference probability of each sample, thereby improving the accuracy of video analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of video analysis technology, and in particular to a video analysis method, device, electronic device, and storage medium. Background Art

[0002] With the development of smart cities and the Internet of Things (IoT), the number of image acquisition devices, such as cameras, has skyrocketed, generating massive amounts of video. Pre-trained video analysis models can be used to identify and locate discrete events in videos. For example, a video analysis model can analyze video features and ultimately output the location of discrete events, such as a high jump, that occurred between minutes X seconds and minutes X seconds.

[0003] In related technologies, when training a video analysis model, the model can be trained based on sample videos annotated with sample labels. To reduce manpower consumption, each sample video can be annotated with a smaller number of sample labels, for example, each sample video can be annotated with one sample label.

[0004] However, precisely because each sample video is annotated with a small number of sample labels, the video analysis model has fewer features to learn. Therefore, the positioning accuracy of the video analysis model trained based on the above method is low, and the accuracy of the positioning information determined based on the video analysis model is also low, that is, the accuracy of the video analysis is low.

[0005] For example, taking the example of a sample video containing a non-continuous event such as high jump, the sample video can be labeled with only the sample label "contains high jump". Under the relevant technology, the video analysis model can identify the most significant segment, such as the moment when the athlete jumps over the pole, and consider that a non-continuous event has been identified. The positioning information output by the video analysis model may only include information about the athlete at the moment of jumping over the pole, while the subject may expect that the positioning information output by the video analysis model can include more information about the athlete during the entire high jump process, for example, more information about the athlete during the entire high jump process, such as from the run-up to jumping over the pole.

[0006] Therefore, there is an urgent need for a technical solution that can improve the accuracy of video analysis. Summary of the Invention

[0007] The embodiments of the present application provide a video analysis method, apparatus, electronic device, and storage medium to improve the accuracy of video analysis.

[0008] In a first aspect, an embodiment of the present application provides a video analysis method, the method comprising:

[0009] Based on the trained target video analysis model, the initial features of each video segment contained in the video to be analyzed are obtained;

[0010] Based on the initial features, obtaining the initial probability that each of the video clips contains each set event;

[0011] Based on the initial probabilities and a preset probability threshold, obtaining location information of the target event contained in the video to be analyzed, and analyzing the target event contained in the video to be analyzed based on the location information;

[0012] Among them, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. The first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is based on the initial probability of each sample, and the sample reference probability that each sample video clip contains each sample event is determined, and is obtained based on each sample reference probability.

[0013] In a second aspect, an embodiment of the present application provides a video analysis device, the device comprising:

[0014] An acquisition module is configured to obtain initial features of each video segment contained in the video to be analyzed based on the trained target video analysis model, and based on each initial feature, obtain an initial probability that each video segment contains each set event;

[0015] A processing module is configured to obtain positioning information of a target event contained in the video to be analyzed based on each initial probability and a preset probability threshold, and analyze the target event contained in the video to be analyzed based on the positioning information;

[0016] Among them, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. The first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is based on the initial probability of each sample, and the sample reference probability that each sample video clip contains each sample event is determined, and is obtained based on each sample reference probability.

[0017] Optionally, the processing module is specifically configured to:

[0018] Based on the video analysis model to be trained, obtaining the sample initial features of each sample video clip and the sample initial probability of each sample video clip containing each sample event;

[0019] Obtaining the first training loss based on the initial probability of each sample and the probability that the sample video contains a label sample of each sample event;

[0020] Determining a sample reference probability that each sample video segment contains each sample event based on a sample initial probability that each sample video segment contains each sample event and a corresponding probability threshold;

[0021] Determining a sample baseline feature of each sample video segment based on the sample initial feature and the corresponding sample reference probability of each sample video segment;

[0022] A second training loss is determined based on the sample baseline features of each sample video clip and a preset sample feature set, wherein the sample feature set includes corresponding positive and negative sample features.

[0023] Optionally, the processing module is specifically configured to:

[0024] Based on the sample video segments, forming sample sub-videos, wherein each sample sub-video includes part of the sample video segments or all of the sample video segments in the sample video segments;

[0025] Determining sample comprehensive features of each sample sub-video based on sample baseline features of the sample video segments contained in each sample sub-video;

[0026] A second training loss is determined based on the sample comprehensive features of each of the sample sub-videos and a preset sample feature set.

[0027] Optionally, the processing module is specifically configured to:

[0028] Determining, based on the sample initial features and the corresponding sample reference probabilities of each sample video clip, a first event sub-feature and a first background sub-feature of each sample video clip;

[0029] The first event sub-feature and the first background sub-feature of each sample video clip are used as sample reference features of each sample video clip.

[0030] Optionally, the processing module is specifically configured to:

[0031] Determining the second event sub-feature of each sample sub-video based on the first event sub-feature of the sample video clip contained in each sample sub-video; and

[0032] determining a second background sub-feature of each sample sub-video based on the first background sub-feature of the sample video segment contained in each sample sub-video;

[0033] The second event sub-feature and the second background sub-feature of each sample sub-video are used as the sample comprehensive features of each sample sub-video.

[0034] Optionally, the processing module is specifically configured to:

[0035] Determining, based on the sample label corresponding to the sample video, the event category of the second event sub-feature of each of the sample sub-videos;

[0036] Determining positive sample features and negative sample features corresponding to each second event sub-feature based on the event category and the category of each sample feature in the sample feature set;

[0037] A second training loss is determined based on a first distance between each second event sub-feature and a corresponding positive sample feature, and a second distance between each second event sub-feature and a corresponding negative sample feature.

[0038] Optionally, the processing module is specifically configured to:

[0039] The following steps are performed for each sample video clip for the sample initial probability of containing each sample event:

[0040] Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event;

[0041] The following operations are performed for each sample video clip: when the sample initial probability that a sample video clip contains the sample event is not less than the probability threshold, the sample reference probability that the sample video clip contains the sample event is configured to a first preset value; otherwise, the sample reference probability that the sample video clip contains the sample event is configured to a second preset value.

[0042] Optionally, the processing module is specifically configured to:

[0043] Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains the sample event according to the value of the sample initial probability that each sample video segment contains the sample event; or

[0044] According to the preset correspondence between each sample event and the probability threshold, the probability threshold corresponding to the sample initial probability that each sample video clip contains the sample event is determined.

[0045] Optionally, the processing module is specifically configured to:

[0046] Determining a corresponding target training loss based on the first training loss, the second training loss, and corresponding weight coefficients, wherein a ratio of the weight coefficient of the first training loss to the weight coefficient of the second training loss is a ratio that is negatively correlated with the total number of rounds of the current iteration;

[0047] According to the target training loss, model parameters of the video analysis model to be trained are adjusted.

[0048] Optionally, the processing module is specifically configured to:

[0049] Performing dimensionality reduction processing on the sample initial features of each sample video clip to obtain sample low-dimensional features corresponding to the sample initial features of each video clip, wherein each sample initial feature has D feature dimensions, each sample low-dimensional feature has d feature dimensions, D>d, and each feature dimension represents a video attribute;

[0050] Based on the low-dimensional features of each sample and the corresponding sample reference probability, a sample baseline feature of each video clip is determined.

[0051] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory, wherein the memory stores a program code, and when the program code is executed by the processor, the processor executes the steps of any one of the above-mentioned video analysis methods.

[0052] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a program code. When the storage medium is run on an electronic device, the program code is used to enable the electronic device to perform the steps of any of the above-mentioned video analysis methods.

[0053] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, the steps of any of the above-mentioned video analysis methods are implemented.

[0054] The beneficial effects of this application are as follows:

[0055] The embodiments of the present application provide a video analysis method, apparatus, electronic device, and storage medium. The target video analysis model of the present application can determine a sample reference probability that each video clip of the sample video contains each sample event based on the sample initial probability that each video clip contains each sample event, and determine a second training loss based on each sample reference probability, and can determine a corresponding training loss for the video analysis model to be trained based on the first training loss and the second training loss.

[0056] Since the present application can determine the sample reference probability that each video clip contains each sample event based on the initial probability of each sample, the sample reference probability can be used as a "pseudo-label" for the electronic device to label each video clip. If the method in which the labeler labels each sample video with a sample label is called video-level labeling, the method in which the electronic device determines the reference probability for each video clip can be called frame-level labeling, compared with the smaller and sparser number of sample labels labeled at the video level, the number of "pseudo-labels" labeled at the frame level is larger and denser. Therefore, compared with the method of weakly supervised training based only on sample labels labeled at the video level in the related art, the present application can perform similar fully supervised training on the video analysis model based on the denser pseudo-labels labeled at the frame level, so that the video analysis model can learn more features. Therefore, the positioning accuracy of the video analysis model trained based on the present application is higher, thereby improving the accuracy of video analysis.

[0057] In addition, since there is no need to increase the number or content of sample labels, the present application can achieve the purpose of improving the positioning accuracy of the video analysis model and thus improving the accuracy of video analysis without increasing manpower consumption.

[0058] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0060] Figure 1 A schematic diagram of an application scenario in an embodiment of the present application;

[0061] Figure 2 A flowchart of an implementation of a video analysis method provided in an embodiment of the present application;

[0062] Figure 3 A flowchart of an implementation method of a video analysis model training method provided in an embodiment of the present application;

[0063] Figure 4 A schematic diagram of a video analysis model training process provided in an embodiment of the present application;

[0064] Figure 5A flowchart for determining the sample reference probability that each sample video clip contains each sample event provided by an embodiment of the present application;

[0065] Figure 6 A flowchart of an implementation of determining sample baseline features for each video clip provided in an embodiment of the present application;

[0066] Figure 7 A schematic diagram of determining the first event sub-feature and the first background sub-feature of each sample video clip provided in an embodiment of the present application;

[0067] Figure 8 A flowchart for determining a second training loss according to an embodiment of the present application is provided;

[0068] Figure 9 A flowchart for determining comprehensive sample features of each sample sub-video provided in an embodiment of the present application;

[0069] Figure 10 A flowchart for determining a second training loss according to an embodiment of the present application is provided;

[0070] Figure 11 A schematic diagram of a feature comparison process provided in an embodiment of the present application;

[0071] Figure 12a A comparison chart showing the effect of different numbers of sample sub-videos on the performance of the video analysis model provided in the embodiment of the present application;

[0072] Figure 12b This is a first experimental comparison effect diagram provided in the embodiment of the present application;

[0073] Figure 12c This is a comparison diagram of the second experiment provided in the embodiment of the present application;

[0074] Figure 12d This is a comparison diagram of the third experiment provided in the embodiment of the present application;

[0075] Figure 13 A schematic diagram of a video analysis process provided in an embodiment of the present application;

[0076] Figure 14 A schematic diagram of a timing flow chart for interactively implementing a video analysis method provided in an embodiment of the present application;

[0077] Figure 15 A schematic diagram of a specific scenario of video analysis provided in an embodiment of the present application;

[0078] Figure 16 A schematic diagram of another specific scenario of video analysis provided in an embodiment of the present application;

[0079] Figure 17 A schematic diagram of the structure of a video analysis device provided in an embodiment of the present application;

[0080] Figure 18 A schematic diagram of the hardware structure of an electronic device to which an embodiment of the present application is applied;

[0081] Figure 19 A schematic diagram of the hardware structure of another electronic device to which an embodiment of the present application is applied. DETAILED DESCRIPTION

[0082] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of the technical solutions of this application, but not all of them. Based on the embodiments described in this application document, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the technical solutions of this application.

[0083] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments of the application described herein can be practiced in orders other than those illustrated or described herein.

[0084] The following explains some of the terms used in the embodiments of the present application to facilitate understanding by those skilled in the art.

[0085] Embedded features: Embedded features can represent an object (such as a word, a product, or a movie) using a feature vector. The property of this feature vector is that objects corresponding to vectors with similar distances have similar meanings. For example, the distance between the Avengers' embedding and the Iron Man's embedding is very close, but the distance between the Avengers' embedding and the Gone with the Wind's embedding is farther. In addition, embeddings can also have mathematical relationships, such as the Madrid embedding - the Spain embedding + the France embedding ≈ the Paris embedding.

[0086] Background: When no set event occurs in the video scene, it can be called background. For example, when there are no players playing in a basketball court, it can be called background.

[0087] Foreground: In contrast to background, when any set event occurs in the video scene, it can be called foreground. For example, when athletes appear on the basketball court playing basketball, it can be called foreground.

[0088] The following is a brief introduction to the design concept of the embodiment of this application:

[0089] When identifying and locating non-continuous events appearing in a video based on a pre-trained video analysis model, for example, the video can be feature analyzed through the video analysis model, and ultimately the location information of non-continuous events such as high jump occurring between the Xth minute and the Xth second and the XXth minute and the XXth second of the video can be output.

[0090] In related technologies, when training a video analysis model, the model can be trained based on sample videos annotated with sample labels. To reduce manpower consumption, each sample video can be annotated with a smaller number of sample labels, for example, each sample video can be annotated with one sample label.

[0091] However, precisely because each sample video is annotated with a small number of sample labels, the video analysis model has fewer features to learn. Therefore, the positioning accuracy of the video analysis model trained based on the above method is low.

[0092] For example, taking the example of a sample video containing a non-continuous event such as high jump, the sample video can be labeled with only the sample label "contains high jump". Under the relevant technology, the video analysis model can identify the most significant segment, such as the moment when the athlete jumps over the pole, and consider that a non-continuous event has been identified. The positioning information output by the video analysis model may only include information about the athlete at the moment of jumping over the pole, while the subject may expect that the positioning information output by the video analysis model can include more information about the athlete during the entire high jump process, for example, more information about the athlete during the entire high jump process, such as from the run-up to jumping over the pole.

[0093] This shows that in related technologies, the positioning accuracy of video analysis models is low.

[0094] In view of this, in order to solve the technical problem of low positioning accuracy of video analysis models in related technologies, the embodiments of the present application propose a video analysis method, device, electronic device and storage medium. The target video analysis model of the present application can determine the sample reference probability that each video clip contains each sample event based on the sample initial probability that each video clip of the sample video contains each sample event, and determine the second training loss based on each sample reference probability, and can determine the corresponding training loss of the video analysis model to be trained based on the first training loss and the second training loss.

[0095] Since the present application can determine the sample reference probability that each video clip contains each sample event based on the initial probability of each sample, the sample reference probability can be used as a "pseudo-label" for the electronic device to mark each video clip. If the method in which the labeler marks a sample label for each sample video is called video-level labeling, the method in which the electronic device determines the reference probability for each video clip can be called frame-level labeling. Compared with the smaller and sparser number of sample labels marked at the video level, the number of "pseudo-labels" marked at the frame level is larger and denser. Therefore, compared with the method of weakly supervised training based only on sample labels marked at the video level in the related art, the present application can perform similar fully supervised training on the video analysis model based on the denser pseudo-labels marked at the frame level, so that the video analysis model can learn more features. Therefore, the positioning accuracy of the video analysis model trained based on the training method of the present application is higher.

[0096] In addition, since there is no need to increase the number or content of sample labels, the present application can achieve the purpose of improving the positioning accuracy of the video analysis model and thus improving the accuracy of video analysis without increasing manpower consumption.

[0097] The preferred embodiments of the present application are described below in conjunction with the drawings in the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application and are not used to limit the present application. In addition, the embodiments and features in the embodiments of the present application can be combined with each other if there is no conflict.

[0098] like Figure 1 , which is a schematic diagram of an application scenario provided by an embodiment of the present application. The application scenario diagram includes two terminal devices 110 and a server 120.

[0099] In the embodiment of the present application, the terminal device 110 includes but is not limited to mobile phones, tablet computers, laptop computers, desktop computers, e-book readers, intelligent voice interaction devices, smart home appliances, car terminals and other devices; a video analysis-related client can be installed on the terminal device, which can be software (such as a browser, video analysis software, etc.), or a web page, applet, etc. The server 120 is a background server corresponding to the software or web page, applet, etc., or a server specifically used for video analysis, which is not specifically limited in this application. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and basic cloud computing services such as big data and artificial intelligence platforms.

[0100] It should be noted that the video analysis method in the embodiment of the present application can be executed by an electronic device, which can be a server 120 or a terminal device 110, that is, the method can be executed by the server 120 or the terminal device 110 alone, or can be executed by the server 120 and the terminal device 110 together. For example, when executed by the terminal device 110 alone, the terminal device 110 can obtain the initial features of each video segment contained in the video to be analyzed based on the trained target video analysis model, and can obtain the initial probability that each video segment contains each set event based on each initial feature; the terminal device 110 can obtain the positioning information of the target event contained in the video to be analyzed based on each initial probability and a preset probability threshold, and analyze the target event contained in the video to be analyzed based on the positioning information; wherein the target video analysis model is obtained by adjusting the parameters of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss, the first training loss is obtained based on the sample initial probability that each sample video segment of the sample video contains each sample event, and the second training loss is determined based on each sample initial probability to determine the sample reference probability that each sample video segment contains each sample event, and is obtained based on each sample reference probability.

[0101] For another example, when the video analysis method in the embodiment of the present application is jointly executed by the server 120 and the terminal device 110, the object can input the video to be analyzed into the terminal device 110 and trigger a video analysis request in the terminal device 110. In response to the video analysis request, the terminal device 110 can send the video analysis request to the server 120. The server 120 can obtain the video to be analyzed and, based on the trained target video analysis model, obtain the initial features of each video clip contained in the video to be analyzed, and based on each initial feature, obtain the initial probability that each video clip contains each set event; the server 120 can obtain the positioning information of the target event contained in the video to be analyzed based on each initial probability and a preset probability threshold, and analyze the target event contained in the video to be analyzed based on the positioning information to obtain a video analysis result. The server 120 can send the video analysis result to the terminal device 110, and the terminal device 110 can receive and present the video analysis result for the object's reference.

[0102] In an optional implementation, the terminal device 110 and the server 120 may communicate via a communication network.

[0103] In an optional implementation, the communication network is a wired network or a wireless network.

[0104] It should be noted that Figure 1The examples shown are just for illustration. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of this application.

[0105] In an embodiment of the present application, when there are multiple servers, the multiple servers can be combined into a blockchain, and the servers are nodes on the blockchain; as in the video analysis method disclosed in the embodiment of the present application, the sample video set, sample feature set and other data involved therein can be stored on the blockchain.

[0106] In addition, the embodiments of the present application can be applied to various scenarios, including not only video analysis scenarios such as video monitoring, video summary generation, and wonderful video detection, but also including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving and other scenarios.

[0107] The following describes the video analysis model training method provided by the exemplary embodiment of the present application in combination with the application scenarios described above and with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown to facilitate understanding of the spirit and principles of the present application, and the implementation of the present application is not limited in this respect.

[0108] It is worth noting that the acquisition, use, and processing of data in the technical solution of this application are in compliance with the relevant provisions of national laws and regulations.

[0109] See Figure 2 As shown, it is a flow chart of an implementation of a video analysis method provided in an embodiment of the present application. The specific implementation process of the method is as follows:

[0110] S21: Based on the trained target video analysis model, obtain initial features of each video segment contained in the video to be analyzed, and based on each initial feature, obtain initial probabilities that each video segment contains each set event.

[0111] In one possible implementation, an object can trigger a preset video analysis request in the electronic device. For example, the object can trigger the video analysis request by clicking a video analysis button, and the electronic device can receive the video analysis request. This application does not specifically limit the method for triggering the video analysis request.

[0112] After receiving a video analysis request, the electronic device can respond to the video analysis request and obtain the corresponding video to be analyzed. Optionally, the video analysis request can include identification information of the video to be analyzed, and the electronic device can obtain the corresponding video to be analyzed based on the identification information. The video identification information can be a video name identifier, a video storage address identifier, etc., and can be flexibly set according to needs. This application does not specifically limit this.

[0113] In one possible implementation, the electronic device may first obtain each video segment contained in the video to be analyzed, and then input each video segment into a target video analysis model. The target video analysis model then performs feature analysis on each video segment to obtain initial features for each video segment contained in the video to be analyzed. Based on each initial feature, the target video analysis model may also obtain an initial probability that each video segment contains each predetermined event. The predetermined events may include any sporting activity such as high jump, diving, soccer, basketball, or volleyball, as well as safety-related events such as the appearance of pedestrians on a road under repair, or highlights from a video. This application does not specifically limit the predetermined events.

[0114] In addition, considering that the process of obtaining each initial feature of the target video analysis model is similar to the process of obtaining each initial feature of the sample by the video analysis model to be trained, we will not go into details here. After introducing the initial features of the sample, we will introduce the process of obtaining each initial feature of the target video analysis model. Similarly, considering that the process of obtaining each initial probability of the target video analysis model is similar to the process of obtaining each initial probability of the sample by the video analysis model to be trained, we will not go into details here. After introducing the initial probabilities of the sample, we will introduce the process of obtaining each initial probability of the target video analysis model.

[0115] S22: Based on the initial probabilities and a preset probability threshold, obtaining location information of the target event contained in the video to be analyzed, and analyzing the target event contained in the video to be analyzed based on the location information;

[0116] Among them, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. The first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is based on the initial probability of each sample, and the sample reference probability that each sample video clip contains each sample event is determined, and is obtained based on each sample reference probability.

[0117] Among them, the process of the target video analysis model obtaining the positioning information will not be described here in detail, and will be introduced after the training process of the video analysis model is introduced later. Optionally, after obtaining the positioning information, the electronic device can analyze the target event contained in the video to be analyzed based on the positioning information to obtain a video analysis result. Exemplarily, the electronic device can directly output the positioning information as the video analysis result of the video to be analyzed for reference by the object, etc. In addition, the electronic device can also edit the video to be analyzed based on the positioning information output by the target video analysis model, edit the part of the video where the target event appears as the video analysis result, and output it for reference by the object, etc. This application does not make specific limitations on this.

[0118] In a possible implementation, the target video analysis model can be obtained by adjusting the parameters of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. Next, the training process of the video analysis model is described in detail. Figure 3 As shown in FIG, it is an implementation flow chart of a video analysis model training method provided in an embodiment of the present application. The specific implementation process of the method is as follows:

[0119] S31: Obtain a set of sample videos, each sample video is set with a corresponding sample label, and each sample label carries the label sample probability of each sample event contained in the corresponding sample video.

[0120] When training a video analysis model to be trained, an electronic device can first obtain a sample video set, which can contain several sample videos. This application does not specifically limit the number of sample videos contained in the sample video set, and it can be flexibly set according to needs.

[0121] Optionally, the electronic device can obtain each sample video from a storage device, download each sample video from a network resource, or generate virtual sample videos based on video standards, etc. This application does not specifically limit the source of the sample videos.

[0122] Optionally, to identify the events contained in the sample videos, each sample video can be set (labeled) with a corresponding sample label. The sample label can carry the label sample probability that the corresponding sample video contains each set event (for ease of description, the probability that the sample video contains the set event carried in the sample label is referred to as the label sample probability). Optionally, to reduce the manpower consumption of labeling sample labels, each sample video can be set with a smaller number of sample labels, for example, each sample video can be set with one sample label, etc.

[0123] In addition, this application does not impose specific restrictions on the content and number of sample events contained in the sample tags (for ease of description, the set events contained in the sample tags are referred to as sample events), and can be flexibly configured according to needs. For example, sample events can include any sporting activity such as high jump, diving, playing football, playing basketball, playing volleyball, etc., as well as safety-related events such as the appearance of pedestrians on a road under maintenance, and even events such as highlights in a video.

[0124] In addition, the present application does not make any specific restrictions on the label sample probability of each sample event contained in the sample label, and can be flexibly set according to needs. For example, the label sample probability can be any value not less than 0 and not greater than 1. Exemplarily, taking the sample events of high jump and kicking a football as examples, the label sample labels corresponding to the sample video can be: the label sample probability of high jump is 1, the label sample probability of kicking a football is 0, etc. Optionally, the labeling personnel can only mark the target events contained (appearing) in the sample video (for the convenience of description, the events appearing in the video are referred to as target events), and the electronic device can configure the label sample probability corresponding to the target event in the sample label to a larger preset value such as 1, and for the sample events not marked by the labeling personnel, the electronic device can configure the label sample probability of the sample event in the sample label to a smaller preset value such as 0, so that the labeling personnel do not need to configure the label sample probability of each sample event in the sample label, which can further save the manpower consumption of the labeling personnel.

[0125] S32: Based on the sample video set, perform at least one round of iterative training on the video analysis model to be trained, and output a corresponding target video analysis model; wherein, in each round of iteration, at least steps S33 to S35 are executed.

[0126] In one possible implementation, after obtaining a sample video set, the electronic device can perform at least one round of iterative training on the video analysis model to be trained based on each sample video in the sample video set, and then output the corresponding trained target video analysis model (for the convenience of description, the trained video analysis model is referred to as the target video analysis model).

[0127] During each round of iterative training, the electronic device may train the video analysis model to be trained based on at least one sample video and adjust the model parameters of the video analysis model to be trained. The electronic device may output the trained target video analysis model after the video analysis model to be trained is trained based on all sample videos in the sample video set. Alternatively, the electronic device may output the trained target video analysis model upon determining that the training loss of the video analysis model to be trained meets the training target, etc., without further limitation.

[0128] The following describes the process of iteratively training a video analysis model based on one sample video. The process of iteratively training a video analysis model based on other sample videos is similar and will not be repeated here. Specifically, the process of iteratively training a video analysis model based on one sample video can be referred to S33 to S35.

[0129] S33: Based on the video analysis model to be trained, obtain the sample initial probability that each sample video segment contained in the sample video contains each sample event, and obtain a first training loss based on each sample initial probability.

[0130] Optionally, based on the video analysis model to be trained, feature analysis can be performed on each sample video segment contained in the sample video to obtain the sample initial features of each sample video segment and the sample initial probability of each sample event contained in each sample segment, and based on the obtained sample initial probabilities and the label sample probability of the sample video containing each sample event, the first training loss is obtained.

[0131] See Figure 4 , which is a schematic diagram of a video analysis model training process provided in an embodiment of the present application. When training the video analysis model, the electronic device can first divide the extracted sample video into multiple sample video segments, and input each divided sample video segment into the video analysis model to be trained. The video analysis model can perform feature analysis on each sample video segment separately to obtain the initial sample features of each sample video segment.

[0132] This application does not impose a specific limit on the number of sample video segments into which the sample video is divided, nor does it impose a specific limit on the number of video frames contained in each sample video segment, which can be flexibly set according to needs. For example, the sample video can be divided into 700 sample video segments, each of which can contain 16 consecutive video frames.

[0133] For ease of understanding, the process of dividing sample video segments provided by this application is illustrated below using a specific embodiment. For example, for a sample video, a video frame can be extracted from the sample video at set time intervals, i.e., the sample video can be divided into frames to obtain each video frame contained in the sample video. Optionally, after the sample video is divided into frames, each video frame obtained can be a video frame in RGB mode.

[0134] In one possible implementation, in order to accurately identify the target event contained in the sample video, after obtaining each video frame contained in the sample video, the optical flow (optical flow between two adjacent video frames, for convenience of description, the optical flow between two video frames is referred to as an optical flow image) between each two adjacent video frames can be determined. The optical flow image can include motion information of entities such as observed objects in the video frame.

[0135] Exemplarily, when acquiring an optical flow image, the following operations may be performed for each video frame: a video frame and the previous adjacent video frame of the video frame are combined into a video frame unit, and an optical flow image between the two adjacent video frames is obtained based on an optical flow algorithm such as the TV-L1 optical flow algorithm. In one possible embodiment, if the total number of video frames contained in the sample video is represented by P, the number of optical flow images determined based on each two adjacent video frames is P-1. For ease of calculation, the total number of optical flow images may be kept consistent with the total number of video frames. Optionally, since the first video frame in the sample video does not have a previous adjacent video frame, in order to keep the total number of optical flow images consistent with the total number of video frames, the first video frame and the first video frame itself may be combined into a video frame unit, and an optical flow image between the two video frames in the video frame unit is obtained based on the optical flow algorithm, so that the total number of optical flow images obtained is also P.

[0136] As another example, when acquiring an optical flow image, the following operations may be performed for each video frame: a video frame and its subsequent adjacent video frame are combined into a video frame group, and an optical flow image between the two adjacent video frames is obtained based on an optical flow algorithm such as the TV-L1 optical flow algorithm. In one possible implementation, in order to keep the total number of optical flow images consistent with the total number of video frames, optionally, since the last video frame in the sample video has no subsequent adjacent video frame, the last video frame and the last video frame itself may be combined into a video frame unit, and an optical flow image between the two video frames in the video frame unit is obtained based on the optical flow algorithm, thereby ensuring that the number of obtained optical flow images is consistent with the number of video frames.

[0137] In one possible implementation, after obtaining each video frame and optical flow image in a sample video, a set number of consecutive video frames (referred to as a first set number for ease of description) and a first set number of optical flow images corresponding to the first set number of video frames can be combined to form a sample video segment. Taking the first set number as 16 as an example, for example, the 1st to 16th video frames and the 1st to 16th optical flow images can be combined to form the first sample video segment, the 17th to 32nd video frames and the 17th to 32nd optical flow images can be combined to form the second sample video segment, and so on.

[0138] In one possible implementation, if the sample video is long, the number of sample video segments obtained may be large, and if the sample video is short, the number of sample video segments obtained may be small. Considering that if the number of sample video segments for feature analysis of the video analysis model to be trained differs greatly for different sample videos, it may affect the positioning accuracy of the video analysis model. Optionally, a set number (referred to as the second set number for the convenience of description) of sample video segments can be extracted for each sample video. The second set number can be flexibly set according to demand. For example, the second set number is 700. For example, when a sample video is long, each video segment consisting of a first set number of video frames and a first set number of optical flow images can be first used as a candidate video segment, and then a second set number of video segments can be selected from the candidate video segments, and the selected second set number of video segments can be used as sample video segments for feature analysis. When a sample video is short, a video segment consisting of a first set number of video frames and a first set number of optical flow images can be first selected as a candidate video segment, and then part of the video segment or all of the video segment in the candidate video segment can be copied at least once and added to the candidate video segment, so that the total number of the added video segments reaches a second set number. The second set number of added video segments can be used as sample video segments for feature analysis.

[0139] In one possible implementation, a second set number of sample video clips can be input into the video analysis model respectively, and the feature extractor in the video analysis model can extract features from each sample video clip one by one. For each sample video clip, a feature vector of each sample video clip can be generated respectively, wherein the feature vector of each sample video clip obtained based on the feature extractor can be considered as a feature obtained based on only a single sample video clip without considering adjacent sample video clips. For the convenience of description, the feature vector of each sample video clip obtained based on the feature extractor is referred to as a single video clip feature (also referred to as Extracted features). The feature dimension of the single video clip feature of each sample video clip can be D, and D can be any positive integer greater than 1. This application does not make specific restrictions on this. If the second set number of sample video clips is represented by T, and the feature dimension of the single video clip feature of each sample video clip is D, then after the sample video is processed by the feature extractor, the feature dimensions obtained are T*D in total.

[0140] In one possible implementation, after obtaining the single video segment features of each sample video segment, the single video segment features of each sample video segment can be input into a temporal convolutional network (temporal conv) in a video analysis model. Optionally, the temporal convolutional network can be a network comprising two layers of convolutional neural networks. Based on the temporal convolutional network, feature analysis is performed on each sample video segment to obtain the sample initial features (also referred to as embedded features) of each sample video segment. Compared to the single video segment features, the sample initial features of each sample video segment can be a feature vector obtained by fusing the single video segment features of multiple adjacent sample video segments.

[0141] Exemplarily, when obtaining the sample initial features of each sample video clip, the following operations can be performed for each sample video clip: a set number of consecutive sample video clips (for the convenience of description, referred to as the third set number) including one sample video clip are respectively regarded as a sample video clip group, the single video clip features of each sample video clip in the sample video clip group are fused, and the fused features are used as the sample initial features (embedded features) of this sample video clip. For example, taking the third set number of 3 as an example, for one of the sample video clips, the sample video clip before the sample video clip, the sample video clip, and the adjacent sample video clip after the sample video clip can be taken as a sample video clip group, and the single video clip features of each sample video clip in the sample video clip group are fused, and the fused features are used as the sample initial features of this sample video clip; in addition, the sample video clip and the two adjacent sample video clips after the sample video clip can be taken as a sample video clip group, and the single video clip features of each sample video clip in the sample video clip group are fused, and the fused features are used as the sample initial features of this sample video clip; in addition, the sample video clip and the two adjacent sample video clips before the sample video clip can be taken as a sample video clip group, and the single video clip features of each sample video clip in the sample video clip group are fused, and the fused features are used as the sample initial features of this sample video clip.

[0142] In one possible implementation, the feature dimension of the sample initial features may be the same as the feature dimension of the single video segment features. For example, if the feature dimension of the single video segment features of each video segment is D, then the feature dimension of the obtained sample initial features of each video segment may also be D. If the second set number of sample video segments is represented by T, then after the sample video is processed by the temporal convolutional network, the feature dimensions of the obtained sample initial features are T*D in total.

[0143] After obtaining the sample initial features of each sample video clip, the sample initial features of each sample video clip can be respectively input into the classification head network in the video analysis model. Based on the classification head network, the sample initial probability that each sample video clip contains each sample event is obtained (for the convenience of description, the probability that each sample video clip contains each sample event obtained based on the classification head network is called the sample initial probability).

[0144] In one possible implementation, after obtaining the sample initial probability that each sample video clip contains each sample event, a first training loss can be obtained based on the obtained sample initial probability and the label sample probability of the sample video (for the convenience of description, the training loss obtained based on the sample initial probability and the label sample probability is referred to as the first training loss).

[0145] Optionally, based on the obtained initial probabilities of each sample and the probability of the sample video containing the label sample of each sample event, the process of obtaining the first training loss can be as follows:

[0146] In one possible implementation, after the sample initial features of each sample video clip are input into the classification head network in the video analysis model, based on the classification head network, in addition to obtaining the sample initial probability that each sample video clip contains each sample event, it is also possible to obtain for each sample video clip the sample probability that each sample video clip does not contain any sample event, that is, each sample video clip is the background (for the convenience of description, the event not contained is referred to as the background). Exemplarily, if the number of sample events is represented by C, then for each sample video clip, C sample initial probabilities of containing each sample event and 1 sample probability of being the background can be obtained, for a total of C+1 probabilities. For the convenience of description, the C sample initial probabilities of containing each sample event and 1 sample probability of being the background for each sample video clip obtained after the classification head network can also be called class activation sequences, and the C sample initial probabilities of containing each sample event and 1 sample probability of being the background for each sample video clip are represented by sample A. b Indicates that it is called C+1 samples A b .

[0147] In one possible implementation, in addition to inputting the single-segment features of each sample video segment obtained by the feature extractor into the temporal convolutional network, the single-segment features of each sample video segment can also be input into the foreground selection network in the video analysis model. In contrast to the background, if the segment that does not contain any sample event is called the background, the segment that contains any sample event can be called the foreground. The foreground selection network can also be called the event selection network.

[0148] The foreground selection network can perform feature analysis on each sample video clip to obtain the sample probability that each sample video clip is a sample video clip containing any sample event (for the convenience of description, the sample probability that each sample video clip obtained based on the foreground selection network is a sample video clip containing any sample event is called the sample foreground probability, foreground score). Each sample video clip corresponds to a sample foreground probability. If the sample foreground probability is represented by Q, assuming that there are T sample video clips, a total of T*1 sample foreground probabilities can be obtained. Among them, the present application does not impose specific restrictions on the specific values of the sample foreground probabilities, and can be flexibly set according to needs. For example, if a sample video clip is identified as containing any sample event, the sample foreground probability of the sample video clip can be configured to a higher probability value such as 1. If a sample video clip is identified as not containing any sample event, that is, when a sample video clip is identified as background, the sample foreground probability of the sample video clip can be configured to a lower probability value such as 0.

[0149] In one possible implementation, the foreground selection network may include two basic fully connected neural network layers, and the probability value output by the foreground selection network may be converted into a foreground probability value within the range of 0 to 1 using a sigmoid function. The sigmoid function may map a real number to the range of 0 to 1 and output a value within the range of 0 to 1. The sigmoid function may be used for binary classification. The formula for the sigmoid function is described as: The value of σ(x) is in the range of 0 to 1.

[0150] For each sample video clip, we obtain the sample foreground probability Q and C+1 sample A for each sample video clip. b Then, the following operations can be performed for each sample video segment: the sample foreground probability Q of any sample video segment is combined with the C+1 sample A of the sample video segment. b The values are multiplied by each other to obtain C+1 updated probability values of the sample video segment. For the convenience of description, the C+1 updated probability values of each sample video segment are represented by sample A. f Indicates that sample A f It is called the sample induction probability.

[0151] For example, assuming that for a sample video clip, the sample events are high jump and playing basketball, the sample A obtained by the classification head network is bThe initial probability of the sample containing high jump is 0.8, the initial probability of the sample containing basketball is 0.2, the probability of the sample being background is 0.1, and the sample foreground probability Q of the sample video clip determined by the foreground selection network is 0.9. Then 0.9 can be multiplied by 0.8, 0.2, and 0.1 respectively, and the final sample A is obtained. f Among the sample events contained in the sample video clip, the probability of including high jump is 0.72, the probability of including playing basketball is 0.18, and the sample video clip does not contain any sample event, and the probability of being background is 0.09.

[0152] Determine C+1 samples A for each sample video clip f and C+1 samples A b Then, we can use Multiple Instance Learning (MIL) technology to determine the C+1 probability that the sample video contains each sample event and the background, based on the C+1 probability values of each sample video segment in the sample video. For ease of description, the sample events and the background are collectively referred to as C+1 categories below.

[0153] Specifically, first take sample A f For example, if the number of sample video clips is represented by T, each sample video clip corresponds to C+1 sample A f , for C+1 categories, the following operations can be performed: from T sample video clips containing a total of T samples A of any category f Among them, select the set number (for the convenience of description, called the fourth set number) of samples A with the highest probability value f , the fourth set number of samples A is selected f The average value of the above is determined as the first estimated total probability that the entire sample video contains this category (for the convenience of description, the probability that each video segment contains a sample A of a category is used as the basis for the prediction of the probability of the entire sample video containing this category). f , the probability that the entire sample video contains this category is called the first estimated total probability). Among them, the fourth set number can be flexibly set according to needs. For example, the fourth set number is represented by K, and K can be T / 8, etc. For the convenience of description, the fourth set number of samples A selected f The average value of the whole sample video is used to determine the first estimated total probability of this category. f Indicates. A total of C+1 first estimated total probabilities P are obtained f .

[0154] For example, taking the sample event of high jump as an example, assuming that 700 sample video clips contain 700 sample A of high jump fAfter sorting in descending order, select 88 samples A with higher probability values f , assuming that these 88 samples A f The average value of is 0.8, then 0.8 can be determined as the first estimated total probability P of the entire sample video containing the sample event of high jump. f .

[0155] Optionally, after determining that the entire sample video contains the first estimated total probability of each category, the first estimated total probability of each category can be normalized by a normalized exponential function (Softmax function) so that the sum of the first estimated total probability of each category is 1.

[0156] In addition, sample A b For example, if the number of sample video clips is represented by T, each sample video clip corresponds to C+1 A b , for C+1 categories, the following operations can be performed separately: from T sample video clips containing a total of T samples A of any category b Among them, select the fourth set number of samples A with the highest probability value b , the fourth set number of samples A is selected b The average value of the above is determined as the second estimated total probability that the entire sample video contains this category (for the convenience of description, the probability that each video segment contains a sample A of a category is used as the basis for the second estimated total probability that the entire sample video contains this category). b , the probability that the entire sample video contains this category is called the second estimated total probability). For the convenience of description, the fourth set number of samples A selected b The second estimated total probability of the entire sample video containing this category is determined by P b Indicates. A total of C+1 first estimated total probabilities P are obtained b .

[0157] For example, taking the background as an example, assuming that the 700 sample video clips are 700 sample A of the background b After sorting in descending order, select 88 samples A with higher probability values b , assuming that these 88 samples A b The average value of is 0.6, then 0.6 can be determined as the second estimated total probability P of the entire sample video containing background b .

[0158] Optionally, after determining that the entire sample video contains the second estimated total probability of each category, the second estimated total probability of each category can be normalized by a normalized exponential function (Softmax function) so that the sum of the second estimated total probability of each category is 1.

[0159] In a possible implementation, the first estimated total probability and the second estimated total probability of the sample video can be determined based on the sample label of the sample video. Optionally, each sample video corresponds to C+1 first estimated total probabilities P f , and C+1 second estimated total probability P b Therefore, the first estimated total probability P f The corresponding first standard total probability (for the convenience of description, the first estimated total probability P f The corresponding standard total probability is called the first standard total probability, and is expressed as y f (denoted by) can also have C+1 sub-probability values. Similarly, the second estimated total probability P f The corresponding second standard total probability (for the convenience of description, the second estimated total probability P b The corresponding standard total probability is called the second standard total probability, and is expressed as y b ) can also have C+1 sub-probability values.

[0160] Among them, whether the first standard total probability y f Or the second standard total probability y b , where the C sub-probability values for the set event can be configured according to the label sample probability of the sample video containing each set event carried in the sample label. For example, if the sample labels corresponding to the sample video are: the label sample probability of high jump is 1, and the label sample probability of playing football is 0, then the first standard total probability y f The sub-probability value of high jump is also 1, and the sub-probability value of playing football is also 0. Similarly, the second standard total probability y b The sub-probability value of high jump is also 1, and the sub-probability value of kicking football is also 0.

[0161] The first standard total probability y f and the second standard total probability y b The difference is that the sub-probability values for the background category are different. Among them, since the first estimated total probability P f It is generated based on the probability that the sample video is the foreground. If a sample video is the foreground, the probability that it is the background can be considered to be 0. Therefore, the first estimated total probability P f The corresponding first standard total probability y f The standard sub-probability value (supervisory signal) for the sample video as the background can be 0.

[0162] The second estimated total probability P b It is generated based on the probability of whether the sample video contains background. Since a sample video usually contains some redundant background even if there is a target event, the second estimated total probability P bThe corresponding second standard total probability y b The standard sub-probability value (supervisory signal) for the sample video containing background can be 1.

[0163] For example, if the sample labels corresponding to the sample video are: the label sample probability of high jump is 1, and the label sample probability of playing football is 0, then the first standard total probability y f The sub-probability value of high jump is 1, the sub-probability value of playing football is 0, and the sub-probability value of background is 0; the total probability of the second standard y b The sub-probability value of high jump is also 1, the sub-probability value of kicking football is also 0, and the sub-probability value of background is 1.

[0164] The first estimated total probability P of each category of the sample video is determined f , the second estimated total probability P b , the first standard total probability y f And the second standard total probability y b After that, the first training loss can be calculated based on the following formula:

[0165]

[0166] in, is the first training loss;

[0167] C is the number of categories of sample events, C+1 is the total number of categories of sample events plus the background category, c=1, 2..., C+1;

[0168] represents the total probability of the second standard for either category;

[0169] represents the second estimated total probability of any category;

[0170] represents the total probability of the first standard for any category;

[0171] represents the first estimated total probability of any one category.

[0172] S34: Based on the initial probabilities of the samples, determine the sample reference probabilities that each sample video clip contains each sample event, and obtain a second training loss based on the sample reference probabilities.

[0173] In one possible implementation, when determining the second training loss, the sample reference probability that each sample video clip contains each sample event can be determined based on the sample initial probability and the corresponding probability threshold of each sample video clip containing each sample event, and the sample baseline feature of each video clip can be determined based on the sample initial feature and the corresponding sample reference probability of each video clip; the second training loss can be determined based on the sample baseline feature of each video clip and a preset sample feature set, wherein the sample feature set includes corresponding positive and negative sample features.

[0174] In one possible implementation, in order to improve the accuracy of the video analysis model, the electronic device can determine the sample reference probability that each sample video clip contains each set event based on the initial probability of each sample. The sample reference probability can be considered as a "virtual label" or "pseudo-label" intelligently annotated by the electronic device for each sample video clip. If the method in which the annotator labels each sample video with a sample label is called video-level annotation, the method in which the electronic device determines the sample reference probability for each sample video clip can be called frame-level annotation. Compared with the smaller and sparser number of sample labels annotated at the video level, the number of "pseudo-labels" annotated at the frame level is larger and denser. Therefore, compared with the method of weakly supervised training based on sample labels annotated at the video level in the related art, the present application can perform similar fully supervised training on the video analysis model based on the denser pseudo-labels annotated at the frame level, so that the video analysis model can learn more feature dimensions. Therefore, the accuracy of the video analysis model trained based on the training method of the present application is higher, and the denser labels based on the frame-level annotation in the present application are based on the intelligent annotation of the electronic device, so the accuracy of the video analysis model can be improved without increasing manpower consumption.

[0175] Optionally, when determining the sample reference probability that each sample video segment contains each sample event based on the sample initial probability that each sample video segment contains each sample event, the sample initial probability that each sample video segment contains each sample event can be directly used to determine the sample reference probability that each sample video segment contains each sample event. For example, if the sample initial probability that a sample video segment contains the sample event "diving" is 1, the sample reference probability that the sample video segment contains the sample event "diving" can be directly configured as 1.

[0176] In a possible implementation, it is also possible to Figure 5 The flowchart shown implements the step of determining the sample reference probability that each sample video segment contains each sample event based on the sample initial probability that each sample video segment contains each sample event and the corresponding probability threshold in S34. Figure 5As shown, it is an implementation flow chart of determining the sample reference probability that each sample video clip contains each sample event provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0177] S34-1: performing the following steps for each sample initial probability of each sample video clip containing each sample event: determining a probability threshold corresponding to the sample initial probability of each sample video clip containing one sample event.

[0178] The following first introduces the process of determining the probability threshold corresponding to the sample initial probability of one sample event as an example. The process of determining the probability threshold corresponding to the sample initial probability of other sample events is similar and will not be repeated here.

[0179] In one possible implementation, when determining the probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event, the preset probability threshold can be directly determined as the probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event. The preset probability threshold can be flexibly set based on needs and is not specifically limited in this application. For example, assuming the preset probability threshold is 0.5, the probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event can be determined as 0.5.

[0180] Optionally, in order to quickly and flexibly determine the probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event, the probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event can also be determined based on the numerical value of the sample initial probability that each sample video clip contains a sample event. Exemplarily, the average value of the sample initial probability that each sample video clip contains a sample event can be determined as the probability threshold corresponding to the sample initial probability that each sample video clip contains this sample event. For example, assuming there are 700 sample video clips, and the average value of the sample initial probability that each of these 700 sample video clips contains the sample event of diving is 0.8, then the sample reference probability that these 700 sample video clips contain the sample event of diving can be configured as 0.8.

[0181] In addition, the maximum value or minimum value of the sample initial probability that each sample video clip contains a sample event can also be determined as the probability threshold corresponding to the sample initial probability that each sample video clip contains this sample event. This application does not make any specific restrictions on this.

[0182] In addition, for the convenience of description, the probability threshold corresponding to the sample initial probability of each sample video clip containing a sample event can be expressed as θ c express.

[0183] Since the present application can determine the probability threshold corresponding to the sample initial probability of each sample video clip containing each sample event based on the numerical value of the sample initial probability of each sample event in each sample video clip, compared to setting the probability threshold to a fixed value, since the probability threshold corresponding to the sample initial probability of each sample event determined by the present application can be adaptively and dynamically adjusted as the sample initial probabilities of different sample videos are different, the probability threshold determined by the present application can be more suitable for the sample video, and the sample reference probability (pseudo-label) can be determined more accurately based on the probability threshold, thereby further improving the positioning accuracy of the video analysis model.

[0184] In addition, in order to quickly and flexibly determine the probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event, the sample probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event can also be determined based on the preset correspondence between each sample event and the probability threshold. Specifically, staff members can pre-set the correspondence between each sample event and the probability threshold and save the correspondence in the electronic device. The probability thresholds corresponding to each sample event can be the same or different and can be flexibly set according to needs. The electronic device can determine the probability threshold corresponding to the sample initial probability that each sample video clip contains any one of the sample events based on the preset correspondence between each sample event and the probability threshold.

[0185] For example, assuming that the preset correspondence between each sample event and the probability threshold includes a correspondence between the sample event diving and a probability threshold of 0.7, and a correspondence between the sample event high jump and a probability threshold of 0.6, the electronic device can configure the probability threshold corresponding to the sample initial probability that each sample video clip contains the sample event diving to 0.7, and configure the probability threshold corresponding to the sample initial probability that each sample video clip contains the sample event high jump to 0.6.

[0186] Since the present application can determine the probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event based on the preset correspondence between each sample event and the probability threshold, it can achieve the purpose of quickly and flexibly determining the probability threshold corresponding to the sample initial probability that each sample video clip contains each sample event.

[0187] S34-2: Perform the following operations for each sample video clip: when the sample initial probability that a sample video clip contains a sample event is not less than the corresponding probability threshold, the sample reference probability that the sample video clip contains the sample event is configured to a first preset value; otherwise, the sample reference probability that the sample video clip contains the sample event is configured to a second preset value.

[0188] After determining the corresponding probability threshold for the sample initial probability of each sample video clip containing a sample event, the sample reference probability (pseudo-label) of each sample video clip containing the sample event can be determined for each sample video clip. The following describes the process of determining the sample reference probability of one sample video clip containing the sample event as an example. The process of determining the sample reference probability of other sample video clips containing the sample event is similar and will not be repeated here.

[0189] Optionally, it can be determined whether the sample initial probability that one of the sample video clips contains a sample event is not less than (greater than or equal to) a corresponding probability threshold. When the sample initial probability that the sample video clip contains a sample event is not less than the corresponding probability threshold, the sample reference probability that the sample video clip contains a sample event can be configured as a first preset value. The first preset value can be flexibly set according to demand, for example, the first preset value can be a larger probability value such as 1. In addition, when the sample initial probability that the sample video clip contains a sample event is less than the corresponding probability threshold, the sample reference probability that the sample video clip contains a sample event can be configured as a second preset value. The second preset value can be flexibly set according to demand, for example, the second preset value can be a smaller probability value such as 0.

[0190] For example, if the probability threshold of each sample video clip containing the sample event of high jump is 0.6, and the sample initial probability of one of the sample video clips containing the sample event of high jump is 0.8, then the sample reference probability (pseudo-label) of the sample video clip containing the sample event of high jump can be configured as 1; and if the sample initial probability of another sample video clip containing the sample event of high jump is 0.5, then the sample reference probability (pseudo-label) of the other sample video clip containing the sample event of high jump can be configured as 0.

[0191] In addition, for the convenience of description, the sample reference probability (pseudo label) of each sample video clip containing each sample event is expressed as Indicated by . Where t is the number identifier of the sample video clips, t = 1, 2, ... T, T is the total number of sample video clips, c is the sample event identifier, c = 1, 2 ..., C, C is the total number of sample events. For example, if the identifier of the sample event high jump is 1, then It can represent the sample reference probability (pseudo label) that the first sample video clip contains the sample event of high jump.

[0192] The present application can quickly and accurately determine the sample reference probability that each sample video clip contains each sample event by comparing the size relationship between the initial probability of each sample and the corresponding probability threshold.

[0193] In a possible implementation, after determining the sample reference probability that each sample video clip contains each sample event, the sample baseline feature of each sample video clip can be determined based on the sample initial feature of each sample video clip and the corresponding sample reference probability.

[0194] In a possible implementation, considering that when the sample initial features have more feature dimensions, if the sample baseline features of each sample video clip are determined directly based on the sample initial features with more feature dimensions, the feature dimensions of the determined sample baseline features are usually also more. When the video analysis model is trained based on the sample baseline features with more feature dimensions, it may make the video analysis model difficult to optimize and prone to overfitting. In order to improve training efficiency and prevent overfitting, the sample initial features of each sample video clip can be first subjected to dimensionality reduction processing to obtain sample low-dimensional features corresponding to the sample initial features of each sample video clip. Based on the sample low-dimensional features and the corresponding sample reference probabilities, when determining the sample baseline features of each sample video clip, the feature dimensions of the sample baseline features of each sample video clip can be reduced.

[0195] Among them, the specific way of performing dimensionality reduction processing on the initial features of the sample is not specifically limited in this application, as long as the feature dimension of the low-dimensional features of the sample after dimensionality reduction processing is smaller than the feature dimension of the initial features of the sample. Exemplarily, the feature dimension of the initial features of the sample can be more (for the convenience of description, referred to as D feature dimensions), and the feature dimension of the low-dimensional features of the sample can be less (for the convenience of description, referred to as d feature dimensions). This application does not specifically limit the specific values of D and d. Optionally, D and d can be positive integers, and D is greater than d. In a possible embodiment, d can be a value such as half of D. Exemplarily, the value of D can be 2048 dimensions, that is, the initial features of the sample can be a 2048-dimensional feature vector, and the value of d can be 1024 dimensions, that is, the low-dimensional features of the sample can be a 1024-dimensional feature vector.

[0196] In one possible implementation, whether it is the feature dimension of the initial feature of the sample or the feature dimension of the low-dimensional feature of the sample, each feature dimension can respectively characterize a video attribute. Exemplarily, the video attributes may include the length, width, time information, scene, entity, action, etc. of the video frame. In one possible implementation, the scene may be the background environment in the video frame, such as a gymnasium, grassland, sky, etc. The entity may be an independent individual appearing in the video frame, such as a person, cat or dog, etc., or a football, basketball, stone, guitar, etc. The action may be the dynamic behavior of a scene or entity in the video frame, and the action may be composed of video frame materials composed of at least one video frame. The characteristics of the action feature dimension can be identified by identifying the internal connection between the entity or scene in at least one video frame. The present application does not specifically limit the video attribute represented by each feature dimension, and it can be flexibly set according to needs.

[0197] Since the present application can determine the sample baseline features of each sample video clip based on the low-dimensional features of the samples with low feature dimensions and the corresponding sample reference probabilities, the feature dimensions of the sample baseline features can also be made low. When the subsequent model training process is based on the sample baseline features with low feature dimensions, the training efficiency can be effectively improved and the model overfitting can be effectively prevented.

[0198] In a possible implementation, it is also possible to Figure 6 The flowchart shown implements the steps of determining the sample baseline features of each sample video segment based on the sample initial features of each sample video segment and the corresponding sample reference probability. Figure 6 As shown in FIG, it is a flowchart of an implementation of determining the sample baseline features of each sample video clip provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0199] S601: Based on the sample initial features and the corresponding sample reference probability of each sample video segment, respectively determine the first event sub-feature and the first background sub-feature of each sample video segment.

[0200] In a possible implementation, in order to distinguish the features when the sample video clip contains an event (for the convenience of description, the features when the sample video clip contains an event are referred to as the first event sub-feature) and the features when the sample video clip does not contain an event as the background (for the convenience of description, the features when the sample video clip does not contain an event as the background are referred to as the first background sub-feature), the first event sub-feature and the first background sub-feature of each sample video clip can be determined based on the sample initial features and the corresponding sample reference probability of each sample video clip.

[0201] Optionally, to improve training efficiency and prevent overfitting, each sample's initial features may be first subjected to dimensionality reduction processing to obtain sample low-dimensional features corresponding to the sample initial features of each sample video clip. Based on each sample low-dimensional feature and the corresponding sample reference probability, the first event sub-feature and first background sub-feature of each sample video clip may be determined. For example, since the sample reference probability is the probability that each sample video clip contains each sample event, the product of the sample low-dimensional feature of each sample video clip and the corresponding sample reference probability may be determined as the first event sub-feature of each sample video clip.

[0202] For example, see Figure 7 , which is a schematic diagram of determining the first event sub-feature and the first background sub-feature of each sample video clip provided by an embodiment of the present application. If the sample low-dimensional features of each sample video clip are represented by X t,i Represented by , where t is the number identifier of the sample video clips, t = 1, 2, ... T, T is the total number of sample video clips. i is the number identifier of the feature dimension, i = 1, 2, ... d, d is the total number of feature dimensions of the sample low-dimensional features. The sample reference probability (pseudo-label) of each sample video clip containing each sample event is expressed as Represents that the first event sub-feature of each sample video clip is

[0203] For example, if a sample video clip contains a sample reference probability (pseudo label) of a high jump sample event If is 1, it can be considered that the sample video clip has the sample event of high jump, and the sample video clip is not the background. Then the first event sub-feature of the sample video clip can be the sample low-dimensional feature X of the sample video clip. t,i If a sample video clip contains the sample event of high jump, the sample reference probability (pseudo label) If it is 0, it can be considered that the sample event of high jump does not exist in the sample video clip, the sample video clip is the background, and the first event sub-feature of the sample video clip is 0.

[0204] Optionally, since the sample reference probability is the probability that each sample video clip contains each sample event, the difference between the preset probability full score value (such as 1) and the corresponding sample reference probability can be considered as the probability that each sample video clip does not contain each sample event, that is, is the background. When determining the first background sub-feature of each sample video clip, the difference between the preset probability full score value and the corresponding sample reference probability can be first determined, and then the product of the sample low-dimensional feature of each sample video clip and the corresponding difference can be determined as the first background sub-feature of each sample video clip.

[0205] For example, if the low-dimensional features of each sample video clip are still represented by X t,i Indicates that each sample video clip contains the sample reference probability (pseudo label) of each sample event. Indicates that the preset probability full score is 1, then the first background sub-feature of each sample video clip is

[0206] For example, if a sample video clip contains the sample event of high jump, the sample reference probability (pseudo label) If it is 1, it can be considered that the sample video clip contains the sample event of high jump. The sample video clip is not the background, so the first background sub-feature of the sample video clip can be 0. If a sample video clip contains the sample event of high jump, the sample reference probability (pseudo label) If it is 0, it can be considered that there is no sample event in the sample video clip, and the sample video clip is the background. Then the first background sub-feature of the sample video clip can be the sample low-dimensional feature X of the sample video clip. t,i .

[0207] In a possible implementation, the dimensionality reduction processing may be omitted for the sample initial features, and the first event sub-features and the first background sub-features of each sample video clip may be determined directly based on the sample initial features with more feature dimensions and the corresponding sample reference probabilities. The process of determining the first event sub-features and the first background sub-features based on the sample initial features is similar to the above-mentioned process of determining the first event sub-features and the first background sub-features based on the sample low-dimensional features. For example, the product of the sample initial features of each sample video clip and the corresponding sample reference probability may be determined as the first event sub-feature of each sample video clip, etc., which will not be repeated here.

[0208] S602: Using the first event sub-feature and the first background sub-feature of each sample video clip as sample reference features of each sample video clip.

[0209] After determining the first event sub-feature and the first background sub-feature of each sample video clip, the first event sub-feature and the first background sub-feature of each sample video clip can be used as the sample baseline feature of each sample video clip. That is, for one of the sample video clips, the sample baseline feature of the sample video clip includes the first event sub-feature and the first background sub-feature of the sample video clip.

[0210] Since the present application can determine the first event sub-features and the first background sub-features of each sample video clip based on the sample initial features and the corresponding sample reference probabilities of each sample video clip, based on the first event sub-features and the first background sub-features, when training the video analysis model, it can help the video analysis model to better distinguish the features between the same event categories, the features between different event categories, and the features between events and backgrounds, thereby better improving the positioning accuracy of the video analysis model.

[0211] After determining the sample baseline features of each sample video clip, the second training loss can be determined based on the sample baseline features of each sample video clip and the preset sample feature set (for the convenience of description, the training loss obtained based on the sample baseline features of each sample video clip and the preset sample feature set is referred to as the second training loss).

[0212] In one possible implementation, the following Figure 8 The flowchart shown implements the step of determining the second training loss based on the sample baseline features of each sample video clip and a preset sample feature set. Figure 8 As shown, it is an implementation flow chart of determining the second training loss provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0213] S801: Based on each sample video segment, form each sample sub-video, wherein each sample sub-video includes a part of the sample video segments or all of the sample video segments in each sample video segment.

[0214] In one possible implementation, considering that directly using the sample baseline features of each sample video clip to determine the second training loss may introduce significant noise, and that the video analysis model may not capture larger-granular features such as the overall global features of the sample video when directly using the sample baseline features of each sample video clip to determine the second training loss, thereby potentially affecting the positioning accuracy of the video analysis model, in order to improve the positioning accuracy of the video analysis model, sample sub-videos may be first formed based on each sample video clip. Each sample sub-video may include some or all of the sample video clips in each sample video clip.

[0215] Among them, this application does not specifically limit the number of sample video clips contained in each sample sub-video. For example, multiple (at least two) sample sub-videos can be composed based on each sample video clip, some of which can contain part of the sample video clips in each sample video clip, and one sample sub-video can contain all the sample video clips, etc. The number of sample video clips in the sample sub-video containing part of the sample video clips can be the same or different, and can be flexibly set according to needs. For example, taking the total number of sample video clips as 700 as an example, a total of 5 sample sub-videos can be composed, of which 4 sample sub-videos each contain 175 sample video clips, and 1 sample sub-video contains 700 sample video clips.

[0216] S802: Determine sample comprehensive features of each sample sub-video based on the sample baseline features of the sample video segments contained in each sample sub-video.

[0217] In one possible implementation, after each sample sub-video is determined, the sample comprehensive features of each sample sub-video can be determined based on the sample baseline features of the sample video clips contained in each sample sub-video (for ease of description, the features of the sample sub-video are referred to as sample comprehensive features).

[0218] In one possible implementation, the following Figure 9 The flowchart shown implements the step of determining the sample comprehensive features of each sample sub-video based on the sample baseline features of the sample video segments contained in each sample sub-video in S802. Figure 9 As shown in FIG, it is an implementation flow chart of determining the comprehensive sample features of each sample sub-video provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0219] S802-1: Determine the second event sub-feature of each sample sub-video based on the first event sub-feature of the sample video clip contained in each sample sub-video; and determine the second background sub-feature of each sample sub-video based on the first background sub-feature of the sample video clip contained in each sample sub-video.

[0220] For the convenience of description, the feature of the sample sub-video containing an event can be called the second event sub-feature, and the second event sub-feature is represented by r m Indicates that the feature when the sample sub-video does not contain an event as the background is called the second background sub-feature, and the second background sub-feature is represented by r' m The following first takes the process of determining the second event sub-feature of one sample sub-video as an example to introduce. The process of determining the second event sub-feature of other sample sub-videos is similar and will not be repeated here.

[0221] In a possible implementation, when determining the second event sub-feature of a sample sub-video based on the first event sub-feature of a sample video clip contained in a sample sub-video, the first event sub-feature, the largest first event sub-feature, the smallest first event sub-feature, etc. of any sample video clip contained in the sample sub-video can be determined as the second event sub-feature of the sample sub-video. This application does not make any specific limitations on this.

[0222] Optionally, when determining the second event subfeature of a sample sub-video based on the first event subfeature of a sample video clip contained in a sample sub-video, the average value of the first event subfeatures of the sample video clips contained in the sample sub-video can also be determined as the second event subfeature of the sample sub-video. Exemplarily, the first event subfeature of the sample video clips in the sample sub-video that is not 0 can be first obtained, and then the average value of the first event subfeature that is not 0 can be used as the second event subfeature of the sample sub-video. For example, if a sample sub-video contains 175 sample video clips, of which 75 have a first event subfeature of 0 and 100 have a first event subfeature that is not 0, the average value of the 100 first event subfeatures that are not 0 can be used as the second event subfeature of the sample sub-video.

[0223] Furthermore, the process for determining the second background sub-feature for each sample sub-video is similar to the process for determining the second event sub-feature for each sample sub-video described above. The following describes the process for determining the second background sub-feature for one sample sub-video as an example. The process for determining the second background sub-feature for other sample sub-videos is similar and will not be repeated here.

[0224] In a possible implementation, when determining the second background sub-feature of a sample sub-video based on the first background sub-feature of a sample video segment contained in a sample sub-video, the first background sub-feature, the largest first background sub-feature, the smallest first background sub-feature, etc. of any sample video segment contained in the sample sub-video may be determined as the second background sub-feature of the sample sub-video. This application does not impose any specific limitation on this.

[0225] Optionally, when determining the second background sub-feature of a sample sub-video based on the first background sub-feature of the sample video clips contained in a sample sub-video, the average value of the first background sub-features of the sample video clips contained in the sample sub-video can also be determined as the second background sub-feature of the sample sub-video. Exemplarily, the first background sub-feature of the sample video clips in the sample sub-video whose first background sub-features are not 0 can be first obtained, and then the average value of the first background sub-features that are not 0 can be used as the second background sub-feature of the sample sub-video. For example, if a sample sub-video contains 175 sample video clips, of which 100 sample video clips have a first background sub-feature of 0 and 75 sample video clips have a first background sub-feature of not 0, the average value of the 75 first background sub-features that are not 0 can be used as the second background sub-feature of the sample sub-video.

[0226] S802-2: Use the second event sub-feature and the second background sub-feature of each sample sub-video as the sample comprehensive feature of each sample sub-video.

[0227] After determining the second event sub-features and second background sub-features of each sample sub-video, the second event sub-features and second background sub-features of each sample sub-video can be used as the sample comprehensive features of each sample sub-video. That is to say, for a certain sample sub-video, the sample comprehensive features of the sample sub-video include the second event sub-features and second background sub-features of the sample sub-video.

[0228] Since the present application can determine the second event sub-features and the second background sub-features of the sample sub-video, based on the second event sub-features and the second background sub-features of the sample sub-video, when training the video analysis model, it can help the video analysis model to better distinguish the features between the same event, the features between different events, and the features between the event and the background, thereby better improving the accuracy of the video analysis model.

[0229] S803: Determine a second training loss based on the sample comprehensive features of each sample sub-video and a preset sample feature set.

[0230] In a possible implementation, the second training loss may be determined based on the sample comprehensive features of each sample sub-video and a preset sample feature set. Figure 10 The flowchart shown implements the step S803. Figure 10 As shown, it is an implementation flow chart of determining the second training loss provided by an embodiment of the present application. The specific implementation process of this method is as follows:

[0231] S803-1: Determine the event category of the second event sub-feature of each sample sub-video based on the sample label corresponding to the sample video.

[0232] In one possible implementation, to accurately determine the positive and negative sample features corresponding to each sample sub-video, the event category of the second event sub-feature of each sample sub-video can be determined based on the sample label corresponding to the sample video. Alternatively, the category of the event present in the sample video (for ease of description, the event present in the video is referred to as the target event) can be determined based on the sample label. For example, if the sample labels of the sample video are: the label sample probability of high jump is 1, and the label sample probability of playing football is 0, then the target event present in the sample video can be considered to be high jump.

[0233] After determining the category of the target event in the sample video based on the sample label corresponding to the sample video, the event category of the second event sub-feature of each sample sub-video can be determined as the category of the target event in the sample video. For example, if the target event in the sample video is high jump, the event category of the second event sub-feature of each sample sub-video can be determined as high jump.

[0234] S803-2: Based on the event category of the second event sub-feature of each sample sub-video and the category of each sample feature in the sample feature set, determine the positive sample feature and the negative sample feature corresponding to each second event sub-feature.

[0235] In one possible implementation, to improve the accuracy of the video analysis model, a sample feature set can be preconfigured. Each sample feature in the sample feature set can correspond to a category identifier, and the category of each sample feature in the sample feature set can be determined based on the category identifier. For example, the categories of sample features may include: various sample events, background, etc. This application does not specifically limit the category identifier, and it can be flexibly set according to needs.

[0236] After determining the event category of the second event sub-feature of each sample sub-video, the electronic device can determine the positive sample features and negative sample features corresponding to each second event sub-feature based on the event category of the second event sub-feature of each sample sub-video and the category of each sample feature in the sample feature set.

[0237] In a possible implementation, when determining the positive sample feature corresponding to the second event sub-feature of each sample sub-video, the following steps may be performed for each sample sub-video:

[0238] For the second event sub-feature of a certain sample sub-video in a sample video, the second event sub-features of other sample sub-videos in the sample video except the sample sub-video can be used as positive sample features of the second event sub-feature of the sample sub-video. In addition, the sample features in the sample feature set whose categories are the same as the event categories of the second event sub-feature of the sample sub-video can also be used as positive sample features of the second event sub-feature of the sample sub-video. In other words, the positive sample feature of the second event sub-feature of a certain sample sub-video can be at least one of the following two features:

[0239] The first type is: the second event sub-features of other sample sub-videos in the sample video.

[0240] The second type is: in the sample feature set, the category of the sample feature is the same as the event category of the second event sub-feature of the sample sub-video.

[0241] For example, if the event category of the second event sub-feature of a sample sub-video of a sample video is high jump, since the event category of the second event sub-feature of each sample sub-video of the same sample video is high jump, the second event sub-features of all sample sub-videos in the sample video other than the sample sub-video can be used as positive sample features of the second event sub-feature of the sample sub-video. In addition, the sample features in the sample feature set with the category of high jump can also be used as positive sample features of the second event sub-feature of the sample sub-video.

[0242] In a possible implementation, when determining the negative sample feature corresponding to the second event sub-feature of each sample sub-video, the following steps may be performed for each sample sub-video:

[0243] For the second event sub-feature of a sample sub-video in a sample video, the second background sub-feature of each sample sub-video in the sample video can be used as a negative sample feature of the second event sub-feature of the sample sub-video. Furthermore, sample features in the sample feature set that are classified as background, as well as sample features whose event category is different from that of the second event sub-feature of the sample sub-video, can also be used as negative sample features of the second event sub-feature of the sample sub-video.

[0244] That is, the negative sample feature of the second event sub-feature of a sample sub-video may be at least one of the following three features:

[0245] The first type is: the second background sub-feature of each sample sub-video in the sample video.

[0246] The second type is: sample feature concentration, category and background sample features.

[0247] The third type is: a sample feature in the sample feature set whose event category is different from the event category of the second event sub-feature of the sample sub-video.

[0248] For example, if the event category of the second event sub-feature of a sample sub-video of a sample video is high jump, the second background sub-feature of each sample sub-video in the sample video, including the sample sub-video, can be used as a negative sample feature of the second event sub-feature of the sample sub-video. In addition, the sample features in the sample feature set whose category is background can be used as negative sample features of the second event sub-feature of the sample sub-video. In addition, the sample features in the sample feature set whose event categories are diving, playing basketball, etc., which are different from the event category of the second event sub-feature of the sample sub-video, can be used as negative sample features of the second event sub-feature of the sample sub-video.

[0249] S803-3: Determine a second training loss based on a first distance between each second event sub-feature and the corresponding positive sample feature, and a second distance between each second event sub-feature and the corresponding negative sample feature.

[0250] After determining the positive sample features and negative sample features corresponding to the second event sub-features of each sample sub-video, the second event sub-features of each sample sub-video can be compared with the corresponding positive sample features and negative sample features. Based on the first distance between the second event sub-features of each sample sub-video and the corresponding positive sample features (for convenience of description, the distance between the second event sub-features and the corresponding positive sample features is referred to as the first distance), and the second distance between the second event sub-features of each sample sub-video and the corresponding negative sample features (for convenience of description, the distance between the second event sub-features and the corresponding negative sample features is referred to as the second distance), a second training loss can be determined, so that the model parameters can be adjusted by the target training loss determined based on the second training loss, so that the video analysis model can correctly identify and distinguish the features of the same event category, the features of different event categories, and the features of events and backgrounds, thereby achieving the purpose of improving the positioning accuracy of the video analysis model.

[0251] Optionally, the second training loss can be calculated using the following formula:

[0252]

[0253] Among them, the second training loss is express;

[0254] The total number of sample sub-videos is M+1, where M is any positive integer;

[0255] m=0, 1, 2, ..., M. When m is equal to 0, it represents the first sample sub-video. When m is equal to 1, it represents the second sample sub-video. ...

[0256] r m Represents the second event sub-feature of any sample sub-video;

[0257] p m Represents a set of positive sample features corresponding to the second event sub-feature of any sample sub-video;

[0258] Represents a positive sample feature corresponding to the second event sub-feature of any sample sub-video;

[0259] represents the first distance between the second event sub-feature and the corresponding positive sample feature;

[0260] N m Represents a negative sample feature set corresponding to the second event sub-feature of any sample sub-video;

[0261] Indicates a positive sample feature or a negative sample feature corresponding to the second event sub-feature of any sample sub-video;

[0262] It represents the first distance between the second event sub-feature and the corresponding positive sample feature, or the second distance between the second event sub-feature and the corresponding negative sample feature.

[0263] τ is a hyperparameter of the adjustment factor. The value of τ can be flexibly set according to the needs. For example, τ can be 0.1, etc.

[0264] Optionally, the second event sub-feature and the second background sub-feature may be normalized based on an L2 function. Based on the normalized second event sub-feature and the second background sub-feature, the first distance and the second distance may be determined more accurately.

[0265] See Figure 11 , which is a schematic diagram of a feature comparison process provided by an embodiment of the present application. For example, for a second event sub-feature r of a sample sub-video, m , by taking the second event sub-feature r m Compare the features with the corresponding positive sample features and negative sample features to determine the second training loss, and adjust the model parameters by the target training loss determined based on the second training loss, so as to bring the second event sub-feature r closer. m The distance between the corresponding positive sample feature (also known as intra-class compactness) and the second event sub-feature r mThe distance between the features of the corresponding negative samples (also known as inter-class separation, foreground-background separation) is expected to continuously reduce the second training loss, so that the video analysis model can correctly identify and distinguish the features of the same event category, the features of different event categories, and the features of events and backgrounds, thereby achieving the purpose of improving the positioning accuracy of the video analysis model.

[0266] Since the present application can accurately determine the second training loss based on the first distance between each second event sub-feature and the corresponding positive sample feature, and the second distance between each second event sub-feature and the corresponding negative sample feature, when the parameters of the video analysis model are adjusted based on the second training loss, the video analysis model can better distinguish between the features of the same event category, the features of different event categories, and the features of events and backgrounds, thereby improving the positioning accuracy of the video analysis model.

[0267] Also, see Figure 12a , which is a comparison chart of the effect of different sample sub-video numbers on the performance of the video analysis model provided by the embodiment of this application, wherein the mean average precision (mAP) is an indicator that can be used to evaluate the performance of the model, and mAP@Avg represents the average value of mAP. Figure 12a It can be seen that as the number of sample sub-videos increases from 1 to 5, the performance of the video analysis model continues to improve. This is because as the number of sample sub-videos increases, the features that the video analysis model can learn during the comparative learning process also increase, and the positioning accuracy of the video analysis model continues to improve. However, when the number of sample sub-videos continues to increase from 5 to 10, the performance of the video analysis model actually decreases. This may be because an excessive number of sample sub-videos may introduce some noise interference to the model learning process, which in turn affects the positioning accuracy of the video analysis model. This also shows that directly using the sample baseline features of 700 sample video clips to determine the second training loss will introduce more noise interference to the model learning process.

[0268] Compared with directly using the sample baseline features of each sample video clip to determine the second training loss, when determining the second training loss based on each sample video clip to form each sample sub-video and based on the sample comprehensive features of each sample sub-video and the preset sample feature set, the video analysis model can better capture larger granular features such as the overall global features of the sample video, and can effectively reduce noise interference, thereby further improving the positioning accuracy of the video analysis model.

[0269] S35: Adjust model parameters based on the target training loss corresponding to the first training loss and the second training loss.

[0270] In one possible implementation, after determining the first training loss and the second training loss, for example, the target training loss of the video analysis model to be trained can be determined based on the sum of the first training loss and the second training loss (for the convenience of description, the training loss of the video analysis model finally determined is referred to as the target training loss), and the model parameters of the video analysis model to be trained can be adjusted.

[0271] Optionally, in order to accurately determine the target training loss, the corresponding target training loss can be determined based on the first training loss, the second training loss and the corresponding weight coefficient. In one possible implementation, considering that the error of the second training loss may be large in the early stage of model training, the weight coefficient corresponding to the second training loss can be configured to a smaller value in the early stage of model training. As the number of model training iterations increases, the weight coefficient corresponding to the second training loss can be gradually increased. Optionally, the ratio of the weight coefficient of the first training loss to the weight coefficient of the second training loss can be configured as a ratio that is negatively correlated with the total number of rounds of the current iteration.

[0272] For example,

[0273] Among them, the target training loss is The weight coefficient of the first training loss can be configured as 1, and the weight coefficient β of the second training loss can be configured as a value that changes dynamically and adaptively with the total number of rounds of the current iteration. For example, as the total number of rounds of the current iteration increases, β can be gradually increased from 0.1 to 1000, etc.

[0274] After determining the target training loss, the model parameters of the video analysis model to be trained can be adjusted according to the target training loss.

[0275] In a specific implementation, when adjusting the model parameters of the video analysis model to be trained, a gradient descent algorithm can be used to backpropagate the gradients of the model parameters of the video analysis model, thereby adjusting the model parameters of the video analysis model and training the video analysis model.

[0276] In a possible implementation, the above operation may be performed on each sample video in the sample video set, and when a preset convergence condition is met, it is determined that the video analysis model training is completed.

[0277] The preset convergence condition may be satisfied by the sample video set passing through the video analysis model to be trained, with the number of sample videos correctly identified exceeding a set number, or the number of iterations of training the video analysis model reaching a set maximum number of iterations. These settings may be flexibly adjusted in specific implementations and are not specifically limited here.

[0278] In one possible implementation, when training a video analysis model, the sample videos in the sample video set can be divided into training sample videos and test sample videos. The video analysis model to be trained is first trained based on the training sample videos, and then the reliability of the trained video analysis model is verified based on the test sample videos.

[0279] Since the present application can determine the corresponding target training loss based on the first training loss, the second training loss and the corresponding weight coefficient, and configure the ratio of the weight coefficient of the first training loss to the weight coefficient of the second training loss to be: a ratio that is negatively correlated with the total number of rounds of the current iteration, the corresponding weight coefficients of the first training loss and the second training loss can be intelligently and dynamically adjusted based on the electronic device, which can further improve the positioning accuracy of the video analysis model and improve the model training efficiency.

[0280] In addition, in the embodiments of the present application, comparative experiments are set up based on different model structures. Figure 12b As shown in the figure, it is the first experimental comparison effect diagram provided by the embodiment of this application, where mAP@IoU=q represents the model's mean average precision (mAP) when the intersection over union (IoU) is q. For example, mAP@IoU=0.1 represents the model's mAP when the IoU is 0.1.

[0281] like Figure 12b As shown, compared with the current mainstream video analysis models on the THUMOS14 public test set, such as UntrimNet

[50] , STPN*

[38] , etc., the performance of the video analysis model in the embodiment of the present application has obvious advantages.

[0282] See Figure 12c , which is a second experimental comparison effect diagram provided in the embodiment of the present application, such as Figure 12c As shown, the performance of the video analysis model in the embodiment of the present application has obvious advantages compared with the current mainstream video analysis models on the ActivityNet v1.2 public test set, such as AutoLoc

[44] , W-TALC

[41] , etc.

[0283] In addition, for the convenience of description, the video analysis model trained by adjusting the model parameters based on the first training loss obtained in the above step S33 is called the baseline model, and the video analysis model trained by adjusting the model parameters based on the target training loss corresponding to the first training loss and the second training loss in this application is called the model of this application. In the embodiment of this application, a comparative experiment is set up based on the baseline model and the model of this application, see Figure 12d As shown, it is a third experimental comparison effect diagram provided in the embodiment of the present application, such as Figure 12d It can be seen that compared with the baseline model, the performance of the video analysis model in the embodiment of the present application has obvious advantages.

[0284] The following explains the process of obtaining the initial features of each video clip contained in the video to be analyzed based on the trained target video analysis model, obtaining the initial probability that each video clip contains each set event based on each initial feature, and obtaining the positioning information of the target event contained in the video to be analyzed based on each initial probability and a preset probability threshold.

[0285] See Figure 13 , which is a schematic diagram of a video analysis process provided by an embodiment of the present application, the method includes the following processes:

[0286] In a possible implementation, similar to the process of obtaining the initial probability of samples during the training of a video analysis model, when performing video analysis on a video to be analyzed, the electronic device can first divide the video to be analyzed into multiple video segments, and input each divided video segment into a target video analysis model. The target video analysis model can perform feature analysis on each video segment separately to obtain the initial features of each video segment.

[0287] For example, see Figure 13 , the second set number of video clips can be input into the target video analysis model respectively, and the feature extractor in the target video analysis model can extract features from each video clip one by one. For each video clip, a feature vector of each video clip can be generated respectively, that is, a single video clip feature (also called Extracted features). Among them, the feature dimension of the single video clip feature of each video clip can be D, and D can be any positive integer greater than 1. This application does not make specific restrictions on this. If the second set number of video clips is represented by T, and the feature dimension of the single video clip feature of each video clip is D, then after the video to be analyzed is processed by the feature extractor, the feature dimensions obtained are T*D in total.

[0288] In one possible implementation, after obtaining the single video segment features for each video segment, each of these features can be input into a temporal convolutional network (TCN) within the target video analysis model. Based on this TCN, feature analysis is performed on each video segment to obtain initial features (also known as embedded features) for each video segment. Compared to single video segment features, the initial features for each video segment can be a feature vector derived from the fusion of single video segment features from multiple adjacent video segments.

[0289] In one possible implementation, the feature dimension of the initial features may be the same as the feature dimension of the single video segment features. For example, if the feature dimension of the single video segment features of each video segment is D, then the feature dimension of the initial features obtained for each video segment may also be D. If the second set number of video segments is represented by T, then after the sample video is processed by the temporal convolutional network, the feature dimensions of the initial features obtained are T*D in total.

[0290] After obtaining the initial features of each video clip, the initial features of each video clip can be input into the classification head network in the target video analysis model respectively. Based on the classification head network, the initial probability of each video clip containing each set event is obtained (for the convenience of description, the probability of each video clip containing each set event obtained based on the classification head network is called the initial probability).

[0291] In one possible implementation, after the initial features of each video clip are input into the classification head network in the video analysis model, based on the classification head network, in addition to obtaining the initial probability that each video clip contains each set event, the probability that each video clip does not contain any set event, that is, each video clip is the background (for the convenience of description, the non-event is referred to as the background) can be obtained for each video clip. Exemplarily, if the number of set events is represented by C, then for each video clip, C initial probabilities of containing each set event and 1 probability of being the background can be obtained, for a total of C+1 probabilities. For the convenience of description, the C initial probabilities of containing each set event and 1 probability of being the background for each video clip obtained after the classification head network can also be called class activation sequences, and the C initial probabilities of containing each set event and 1 probability of being the background for each video clip are represented by A. b It is called C+1A b .

[0292] In one possible implementation, in addition to inputting the single-segment features of each video segment obtained by the feature extractor into the temporal convolutional network, the single-segment features of each video segment can also be input into the foreground selection network in the target video analysis model. In contrast to the background, if the segment that does not contain any set event is called the background, the segment that contains any set event can be called the foreground. The foreground selection network can also be called the event selection network.

[0293] The foreground selection network can perform feature analysis on each video clip to obtain the probability that each video clip contains any set event (for the convenience of description, the probability that each video clip obtained based on the foreground selection network contains any set event is called the foreground probability, foreground score). Each video clip corresponds to a foreground probability. If the foreground probability is represented by Q, assuming there are T video clips in total, a total of T*1 foreground probabilities can be obtained.

[0294] For each video clip, we get 1 foreground probability Q and C+1 A for each video clip. b Then, the following operations can be performed for each video clip: the foreground probability Q of any video clip is added to the C+1 A of the video clip. b The values are multiplied respectively to obtain C+1 updated probability values of the video segment. For the convenience of description, the C+1 updated probability values of each video segment are represented by A f Indicates that A f It is called the inductive probability. Among them, each set event and background is collectively referred to as C+1 categories, that is, for each video clip, we can get the A corresponding to each of the C+1 categories. f value.

[0295] Similar to the above embodiment, for example, assuming that for a certain video clip, the events are set as high jump and playing basketball, the A obtained by the classification head network is b In the example, the initial probability of high jump is 0.8, the initial probability of playing basketball is 0.2, the probability of being background is 0.1, and the foreground probability Q of the video clip determined by the foreground selection network is 0.9. Then 0.9 can be multiplied by 0.8, 0.2, and 0.1 respectively, and the final A is obtained. f Among the set events contained in the video clip, the probability of including high jump is 0.72, the probability of including playing basketball is 0.18, and the probability of the video clip not containing any set event and being the background is 0.09.

[0296] Determine the A corresponding to each of the C+1 categories of each video clip fAfter the value is obtained, taking a certain video clip as an example, the A values corresponding to the C+1 categories of the video clip can be calculated. f The values are compared with the preset probability thresholds. If the A f If the value is not less than (greater than or equal to) the preset probability threshold, the video clip can be considered as a video clip of this category. For example, if the preset probability threshold is 0.6, the A value of a video clip is not less than (greater than or equal to) the preset probability threshold. f In the video, the probability of the set event containing the high jump category is 0.72, so the video clip can be considered as a video clip containing high jump. f In the example, another video clip does not contain any set event, that is, the probability that it is the background is 0.8, then it can be considered as a video clip that does not contain any set event, that is, it is a background video clip.

[0297] Optionally, if for a certain type of set event, there is at least one (or more) video clip containing the inductive probability A of the set event f If the probability of each event is not less than a preset probability threshold, the set event can be used as the target event in the video to be analyzed, and the location information of the target event in the video to be analyzed can be determined based on the time information of the at least one (or more) video clips in the video to be analyzed. For example, if the video to be analyzed contains a total of 700 video clips, the A of the set event high jump is included in the 10th to 50th video clips. f are not less than a preset probability threshold, and the 10th video segment is the video segment that starts playing at the 4th minute and 30th second of the video to be analyzed, and the 50th video segment is the video segment that ends at the 6th minute and 20th second of the video to be analyzed, then the positioning information of the target event contained in the video to be analyzed can be: the 4th minute and 30th second to the 6th minute and 20th second of the video is the high jump.

[0298] In one possible implementation, since the video analysis model has determined the second training loss based on the above-mentioned feature comparison process during the training process, and adjusted the model parameters based on the first training loss and the second training loss, the target event contained in the video to be analyzed can be directly identified and located based on the model parameters of the trained target video analysis model, without the need to perform the feature comparison process in the above-mentioned model training process such as step S34. Therefore, compared with the target video analysis model in the related art that is only trained based on the first training loss, the target video analysis model of the present application does not add any additional calculation process. Therefore, the positioning information of the target event contained in the video to be analyzed can be obtained quickly and accurately while ensuring less time consumption.

[0299] In one possible implementation, after obtaining location information of a target event appearing in a video based on a target video analysis model, the electronic device may display this location information for reference. Optionally, the electronic device may also edit the video based on this location information, specifically clipping the portion of the video where the target event appears for reference. Furthermore, the electronic device may save only the clipped portion of the video where the target event appears, while deleting other irrelevant and redundant content in the video to be analyzed, to conserve storage space.

[0300] See Figure 14 As shown, it is a schematic diagram of the interactive implementation timing flow of a video analysis method provided by an embodiment of the present application. The specific implementation process of the method is as follows:

[0301] S1401: The subject inputs the video to be analyzed into the terminal device, and the terminal device receives the video analysis request.

[0302] S1402: The terminal device responds to the video analysis request and sends the video analysis request to the server.

[0303] S1403: The server obtains the video to be analyzed, and performs feature analysis on the video to be analyzed based on the target video analysis model stored in the server, obtains the positioning information of the target event contained in the video, analyzes the target event contained in the video to be analyzed based on the positioning information, obtains the video analysis result, and sends the video analysis result to the terminal device.

[0304] S1404: The terminal device receives the video analysis result and displays it.

[0305] For ease of understanding, the video analysis process provided in this application is illustrated below through a specific embodiment.

[0306] See Figure 15, which is a schematic diagram of a specific scenario of video analysis provided by an embodiment of the present application. Take the video to be analyzed as an example, which contains the entire high jump process of a high jump athlete from the run-up to the overpass. The subject can input the video to be analyzed into the terminal device and enter the event category identifier of the high jump in the input box, such as entering high jump or 1, and trigger the video analysis request by clicking the video analysis button. After the terminal device receives the recommendation request triggered by the object, it can send the recommendation request to the server, wherein the server stores the trained target video analysis model. After receiving the video analysis request, the server can divide the video to be analyzed into 700 video segments in response to the video analysis request, and input the 700 video segments into the target video analysis model in sequence. The target video analysis model performs feature analysis on the 700 video segments in sequence to obtain the positioning information of the target event high jump contained in the video: the 4th minute 30 seconds to the 6th minute 20 seconds of the video is the high jump, and the positioning information contains more information about the athlete's entire high jump process from the run-up to the overpass. Optionally, the server may clip the 4th minute 30 seconds to the 6th minute 20 seconds of the video to be analyzed, and send the clipped shorter video to the terminal device, which displays the clipped video for reference by the subject.

[0307] See Figure 16 , which is a schematic diagram of another specific scenario of video analysis provided by an embodiment of the present application. Take the video to be analyzed that contains the appearance of a pedestrian in a maintenance section as an example. The subject can input the video to be analyzed into the terminal device, and can input the event category identifier of the appearance of a pedestrian in the maintenance section in the input box, and trigger the video analysis request by clicking the video analysis button, etc. The terminal device stores the trained target video analysis model. After the terminal device receives the recommendation request triggered by the object, in response to the video analysis request, the video to be analyzed can be divided into multiple video clips, and the multiple video clips are input into the target video analysis model in sequence. The target video analysis model performs feature analysis on the multiple video clips in sequence to obtain the positioning information of the target event pedestrian appearing in the maintenance section contained in the video: the 10th minute 30 seconds to the 30th minute 50 seconds of the video is the appearance of a pedestrian in the maintenance section, and the positioning information can include the entire process of the pedestrian from the beginning of appearing in the maintenance section to the final departure from the maintenance section. The terminal device can display the video analysis results determined based on the positioning information.

[0308] Based on the same inventive concept, the embodiment of the present application also provides a video analysis device. Figure 17 As shown, it is a structural diagram of a video analysis device 1700 provided in an embodiment of the present application. The device may include:

[0309] Acquisition module 1701: used to obtain initial features of each video segment contained in the video to be analyzed based on the trained target video analysis model, and obtain initial probabilities that each video segment contains each set event based on each initial feature;

[0310] Processing module 1702: configured to obtain location information of a target event contained in the video to be analyzed based on each initial probability and a preset probability threshold, and analyze the target event contained in the video to be analyzed based on the location information;

[0311] Among them, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. The first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is based on the initial probability of each sample, and the sample reference probability that each sample video clip contains each sample event is determined, and is obtained based on each sample reference probability.

[0312] Optionally, the processing module 1702 is specifically configured to:

[0313] Based on the video analysis model to be trained, obtaining the sample initial features of each sample video clip and the sample initial probability of each sample event;

[0314] Obtain a first training loss based on the initial probability of each sample and the probability of the sample video containing the label sample of each sample event;

[0315] Determining a sample reference probability that each sample video clip contains each sample event based on a sample initial probability that each sample video clip contains each sample event and a corresponding probability threshold;

[0316] Determining a sample baseline feature of each sample video segment based on the sample initial feature and the corresponding sample reference probability of each sample video segment;

[0317] A second training loss is determined based on the sample baseline features of each sample video clip and a preset sample feature set, where the sample feature set includes corresponding positive and negative sample features.

[0318] Optionally, the processing module 1702 is specifically configured to:

[0319] Based on each sample video segment, forming each sample sub-video, wherein each sample sub-video contains part of the sample video segment or all of the sample video segment in each sample video segment;

[0320] Determining sample comprehensive features of each sample sub-video based on sample baseline features of the sample video segments contained in each sample sub-video;

[0321] A second training loss is determined based on the sample comprehensive features of each sample sub-video and a preset sample feature set.

[0322] Optionally, the processing module 1702 is specifically configured to:

[0323] Determining, based on the sample initial features and the corresponding sample reference probabilities of each sample video clip, a first event sub-feature and a first background sub-feature of each sample video clip;

[0324] The first event sub-feature and the first background sub-feature of each sample video clip are used as sample reference features of each sample video clip.

[0325] Optionally, the processing module 1702 is specifically configured to:

[0326] Determining the second event sub-feature of each sample sub-video based on the first event sub-feature of the sample video clip contained in each sample sub-video; and

[0327] Determining a second background sub-feature of each sample sub-video based on the first background sub-feature of the sample video clip contained in each sample sub-video;

[0328] The second event sub-feature and the second background sub-feature of each sample sub-video are used as the sample comprehensive features of each sample sub-video.

[0329] Optionally, the processing module 1702 is specifically configured to:

[0330] Determining, based on the sample label corresponding to the sample video, the event category of the second event sub-feature of each sample sub-video;

[0331] Determine the positive sample features and negative sample features corresponding to each second event sub-feature based on the event category and the category of each sample feature in the sample feature set;

[0332] A second training loss is determined based on a first distance between each second event sub-feature and the corresponding positive sample feature, and a second distance between each second event sub-feature and the corresponding negative sample feature.

[0333] Optionally, the processing module 1702 is specifically configured to:

[0334] The following steps are performed for each sample video clip for the sample initial probability of containing each sample event:

[0335] Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event;

[0336] The following operations are performed for each sample video clip: when the sample initial probability that a sample video clip contains a sample event is not less than a probability threshold, the sample reference probability that the sample video clip contains the sample event is configured to a first preset value; otherwise, the sample reference probability that the sample video clip contains the sample event is configured to a second preset value.

[0337] Optionally, the processing module 1702 is specifically configured to:

[0338] Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains the sample event according to the value of the sample initial probability that each sample video segment contains the sample event; or

[0339] According to the preset correspondence between each sample event and the probability threshold, the probability threshold corresponding to the sample initial probability that each sample video clip contains the sample event is determined.

[0340] Optionally, the processing module 1702 is specifically configured to:

[0341] Determine a corresponding target training loss based on the first training loss, the second training loss, and corresponding weight coefficients, wherein a ratio of the weight coefficient of the first training loss to the weight coefficient of the second training loss is a ratio that is negatively correlated with the total number of rounds of the current iteration;

[0342] According to the target training loss, the model parameters of the video analysis model to be trained are adjusted.

[0343] Optionally, the processing module 1702 is specifically configured to:

[0344] Perform dimensionality reduction processing on the sample initial features of each sample video clip to obtain the sample low-dimensional features corresponding to the sample initial features of each video clip, wherein the feature dimensions of each sample initial feature are D, the feature dimensions of each sample low-dimensional feature are d, D>d, and each feature dimension represents a video attribute;

[0345] Based on the low-dimensional features of each sample and the corresponding sample reference probability, the sample baseline features of each video clip are determined.

[0346] For the convenience of description, the above parts are divided into modules (or units) according to their functions and described separately. Of course, when implementing this application, the functions of each module (or unit) can be implemented in the same or multiple software or hardware.

[0347] Those skilled in the art will appreciate that various aspects of the present application can be implemented as systems, methods, or program products. Therefore, various aspects of the present application can be specifically implemented in the following forms: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation that combines hardware and software aspects, which may be collectively referred to herein as a "circuit," "module," or "system."

[0348] Based on the same inventive concept as the above method embodiment, an electronic device is also provided in the embodiment of the present application. In one embodiment, the electronic device may be a server, such as Figure 1 In this embodiment, the structure of the electronic device can be as follows: Figure 18 As shown, it includes a memory 1801 , a communication module 1803 and one or more processors 1802 .

[0349] Memory 1801 is used to store computer programs executed by processor 1802. Memory 1801 may primarily include a program storage area and a data storage area. The program storage area may store an operating system and programs required for running instant messaging functions, while the data storage area may store various instant messaging messages and operating instruction sets.

[0350] Memory 1801 may be a volatile memory, such as random-access memory (RAM); a non-volatile memory, such as read-only memory, flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or any other medium capable of carrying or storing a desired computer program in the form of instructions or data structures and accessible by a computer, but is not limited thereto. Memory 1801 may be a combination of the aforementioned memories.

[0351] The processor 1802 may include one or more central processing units (CPUs) or digital processing units, etc. The processor 1802 is configured to implement the above-mentioned video analysis method when calling the computer program stored in the memory 1801 .

[0352] The communication module 1803 is used to communicate with terminal devices and other servers.

[0353] The specific connection medium between the memory 1801, the communication module 1803 and the processor 1802 is not limited in the embodiment of the present application. Figure 18 In the embodiment, the memory 1801 and the processor 1802 are connected via a bus 1804. Figure 18 The connections between the other components are shown in bold lines for illustration only and are not intended to be limiting. The bus 1804 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 18 The diagram shows a single thick line, but this does not indicate that there is only one bus or one type of bus.

[0354] The memory 1801 stores a computer storage medium, which stores computer executable instructions. The computer executable instructions are used to implement the video analysis method of the embodiment of the present application. The processor 1802 is used to execute the above-mentioned video analysis method, such as Figure 2 shown.

[0355] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 The terminal device 110 shown in FIG. In this embodiment, the structure of the electronic device can be as follows: Figure 19 As shown, it includes: a communication component 1910, a memory 1920, a display unit 1930, a camera 1940, a sensor 1950, an audio circuit 1960, a Bluetooth module 1970, a processor 1980 and other components.

[0356] The communication component 1910 is used to communicate with the server. In some embodiments, it may include a wireless fidelity (WiFi) module. The WiFi module is a short-range wireless transmission technology. Electronic devices can help users send and receive information through the WiFi module.

[0357] The memory 1920 can be used to store software programs and data. The processor 1980 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 1920. The memory 1920 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. The memory 1920 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 1920 can store the operating system and various application programs, and may also store code for executing the video analysis method of the embodiment of the present application.

[0358] The display unit 1930 can also be used to display information input by the user or provided to the user, as well as a graphical user interface (GUI) for displaying various menus of the terminal device 110. Specifically, the display unit 1930 may include a display screen 1932 disposed on the front of the terminal device 110. The display screen 1932 may be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 1930 can be used to display the positioning information interface in the embodiments of the present application, etc.

[0359] The display unit 1930 can also be used to receive input digital or character information and generate signal input related to user settings and function control of the terminal device 110. Specifically, the display unit 1930 may include a touch screen 1931 set on the front of the terminal device 110, which can collect user touch operations on or near it, such as clicking a button, dragging a scroll box, etc.

[0360] The touch screen 1931 can be covered on the display screen 1932, or the touch screen 1931 and the display screen 1932 can be integrated to realize the input and output functions of the terminal device 110. The integrated touch screen can be simply referred to as a touch display screen. In this application, the display unit 1930 can display applications and corresponding operation steps.

[0361] Camera 1940 can be used to capture still images, and users can post comments on images captured by camera 1940 through the application. Camera 1940 can be one or more. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, which is then transmitted to processor 1980 for conversion into a digital image signal.

[0362] The terminal device may further include at least one sensor 1950, such as an acceleration sensor 1951, a distance sensor 1952, a fingerprint sensor 1953, and a temperature sensor 1954. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.

[0363] The audio circuit 1960, speaker 1961, and microphone 1962 provide an audio interface between the user and the terminal device 110. The audio circuit 1960 can convert the received audio data into an electrical signal and transmit it to the speaker 1961, which converts it into a sound signal for output. The terminal device 110 may also be equipped with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 1962 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1960 and converted into audio data. The audio data is then output to the communication component 1910 for transmission to, for example, another terminal device 110, or the audio data is output to the memory 1920 for further processing.

[0364] The Bluetooth module 1970 is used to exchange information with other Bluetooth devices having a Bluetooth module through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 1970 to exchange data.

[0365] Processor 1980 is the control center of the terminal device, connecting various components of the entire terminal using various interfaces and lines. It executes various terminal functions and processes data by running or executing software programs stored in memory 1920 and accessing data stored in memory 1920. In some embodiments, processor 1980 may include one or more processing units. Processor 1980 may also integrate an application processor and a baseband processor, where the application processor primarily handles the operating system, user interface, and application programs, while the baseband processor primarily handles wireless communications. It is understood that the baseband processor may not be integrated into processor 1980. In this application, processor 1980 can run the operating system, application programs, user interface display and touch response, as well as the video analysis method of the embodiments of this application. In addition, processor 1980 is coupled to display unit 1930.

[0366] In some possible implementations, various aspects of the video analysis method provided in the present application may also be implemented in the form of a program product, which includes a computer program. When the program product is run on a computer device, the computer program is used to enable the computer device to execute the steps of the video analysis method according to various exemplary embodiments of the present application described above in this specification. For example, the computer device may execute the following steps: Figure 2 Follow the steps shown in .

[0367] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0368] The program product of the embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and can be run on a computing device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with a command execution system, device, or device.

[0369] A readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries a readable computer program. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with a command execution system, apparatus, or device.

[0370] The computer program embodied on the readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0371] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, and the like, as well as conventional procedural programming languages such as "C" or similar programming languages. The computer program may be executed entirely on the user computing device, partially on the user device, as a stand-alone software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0372] It should be noted that although several units or subunits of the device are mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, depending on the embodiment of the application, the features and functions of two or more units described above can be embodied in a single unit. Conversely, the features and functions of a single unit described above can be further divided and embodied by multiple units.

[0373] Furthermore, although the operations of the method of the present application are described in a particular order in the accompanying drawings, this does not require or imply that the operations must be performed in this particular order, or that all illustrated operations must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.

[0374] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0375] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0376] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0377] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0378] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.

Claims

1. A video analysis method, characterized in that: The method comprises: Based on the trained target video analysis model, obtaining the single video segment features of each video segment contained in the video to be analyzed; Perform feature analysis on the single video segment features of each video segment to obtain the foreground probability of each video segment and the initial features of each video segment; Based on the initial features, obtaining an initial probability that each of the video clips contains each set event and a probability that each of the video clips is background; For each video segment, respectively performing: based on the foreground probability of the video segment, respectively updating the initial probabilities of the video segment and the probability that the video segment is background; Obtaining positioning information of a target event contained in the video to be analyzed based on the updated initial probabilities, the probability that each of the video clips is background, and a preset probability threshold, and analyzing the target event contained in the video to be analyzed based on the positioning information; The target video analysis model is obtained by adjusting parameters of the video analysis model to be trained based on a target training loss corresponding to a first training loss and a second training loss, wherein the first training loss is obtained based on the sample initial probability that each sample video segment of the sample video contains each sample event; The second training loss is obtained as follows: Determining a sample reference probability that each sample video segment contains each sample event based on a sample initial probability that each sample video segment contains each sample event and a corresponding probability threshold; Determining, based on the sample initial features and the corresponding sample reference probabilities of each sample video clip, a first event sub-feature and a first background sub-feature of each sample video clip; Using the first event sub-feature and the first background sub-feature of each sample video clip as the sample baseline feature of each sample video clip; A second training loss is determined based on the sample baseline features of each sample video clip and a preset sample feature set, wherein the sample feature set includes corresponding positive and negative sample features.

2. The method according to claim 1, characterized in that During the training process, the first training loss is obtained in the following way: Based on the video analysis model to be trained, obtaining the sample initial features of each sample video clip and the sample initial probability of each sample video clip containing each sample event; The first training loss is obtained based on the initial probability of each sample and the probability of the label sample of each sample event contained in the sample video.

3. The method according to claim 2, characterized in that The determining of the second training loss based on the sample baseline feature of each sample video clip and a preset sample feature set includes: Based on the sample video segments, forming sample sub-videos, wherein each sample sub-video includes part of the sample video segments or all of the sample video segments in the sample video segments; Determining sample comprehensive features of each sample sub-video based on sample baseline features of the sample video segments contained in each sample sub-video; A second training loss is determined based on the sample comprehensive features of each of the sample sub-videos and a preset sample feature set.

4. The method according to claim 3, characterized in that The determining of the sample comprehensive features of each sample sub-video based on the sample baseline features of the sample video segments contained in each sample sub-video includes: Determining the second event sub-feature of each sample sub-video based on the first event sub-feature of the sample video clip contained in each sample sub-video; and determining a second background sub-feature of each sample sub-video based on the first background sub-feature of the sample video segment contained in each sample sub-video; The second event sub-feature and the second background sub-feature of each sample sub-video are used as the sample comprehensive features of each sample sub-video.

5. The method according to claim 3, characterized in that The determining of the second training loss based on the sample comprehensive features of each of the sample sub-videos and a preset sample feature set includes: Determining, based on the sample label corresponding to the sample video, the event category of the second event sub-feature of each of the sample sub-videos; Determining positive sample features and negative sample features corresponding to each second event sub-feature based on the event category and the category of each sample feature in the sample feature set; A second training loss is determined based on a first distance between each second event sub-feature and a corresponding positive sample feature, and a second distance between each second event sub-feature and a corresponding negative sample feature.

6. The method according to claim 1, characterized in that The determining, based on the sample initial probability that each sample video segment contains each sample event and the corresponding probability threshold, the sample reference probability that each sample video segment contains each sample event, includes: The following steps are performed for each sample video clip for the sample initial probability of containing each sample event: Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains a sample event; The following operations are performed for each sample video clip: when the sample initial probability that a sample video clip contains the sample event is not less than the probability threshold, the sample reference probability that the sample video clip contains the sample event is configured to a first preset value; otherwise, the sample reference probability that the sample video clip contains the sample event is configured to a second preset value.

7. The method according to claim 6, characterized in that The step of determining the probability threshold corresponding to the sample initial probability that each sample video clip contains a sample event includes: Determine a probability threshold corresponding to the sample initial probability that each sample video segment contains the sample event according to the value of the sample initial probability that each sample video segment contains the sample event; or According to the preset correspondence between each sample event and the probability threshold, the probability threshold corresponding to the sample initial probability that each sample video clip contains the sample event is determined.

8. The method according to any one of claims 1-3, 4-7, characterized in that: Based on the target training loss corresponding to the first training loss and the second training loss, parameters of the video analysis model to be trained are adjusted, including: Determining a corresponding target training loss based on the first training loss, the second training loss, and corresponding weight coefficients, wherein a ratio of the weight coefficient of the first training loss to the weight coefficient of the second training loss is a ratio that is negatively correlated with the total number of rounds of the current iteration; According to the target training loss, model parameters of the video analysis model to be trained are adjusted.

9. The method according to any one of claims 2-3, 4-7, characterized in that: The determining, based on the sample initial features and the corresponding sample reference probabilities of each sample video segment, respectively, the first event sub-feature and the first background sub-feature of each sample video segment includes: Performing dimensionality reduction processing on the sample initial features of each sample video clip to obtain sample low-dimensional features corresponding to the sample initial features of each video clip, wherein each sample initial feature has D feature dimensions, each sample low-dimensional feature has d feature dimensions, D>d, and each feature dimension represents a video attribute; Based on the low-dimensional features of each sample and the corresponding sample reference probability, a first event sub-feature and a first background sub-feature of each sample video clip are determined respectively.

10. A video analysis device, characterized in that: The device comprises: The acquisition module is configured to obtain, based on the trained target video analysis model, the single video segment features of each video segment contained in the video to be analyzed; perform feature analysis on the single video segment features of each video segment to obtain the foreground probability of each video segment and the initial features of each video segment; and based on the initial features, obtain the initial probability that each video segment contains each set event and the probability that each video segment is the background; A processing module is configured to, for each video segment, respectively perform the following operations: updating each initial probability of the video segment and the probability of the video segment being the background based on the foreground probability of the video segment; obtaining positioning information of a target event contained in the video to be analyzed based on the updated initial probabilities, the probability of each video segment being the background, and a preset probability threshold; and analyzing the target event contained in the video to be analyzed based on the positioning information; In which, the target video analysis model is obtained after parameter adjustment of the video analysis model to be trained based on the target training loss corresponding to the first training loss and the second training loss. The first training loss is obtained based on the sample initial probability that each sample video clip of the sample video contains each sample event, and the second training loss is obtained in the following manner: based on the sample initial probability that each sample video clip contains each sample event and the corresponding probability threshold, determining the sample reference probability that each sample video clip contains each sample event; based on the sample initial features and the corresponding sample reference probability of each sample video clip, respectively determining the first event sub-feature and the first background sub-feature of each sample video clip; using the first event sub-feature and the first background sub-feature of each sample video clip as the sample baseline features of each sample video clip; determining the second training loss based on the sample baseline features of each sample video clip and a preset sample feature set, wherein the sample feature set contains corresponding positive and negative sample features.

11. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores program codes, and when the program codes are executed by the processor, the processor executes the steps of the method according to any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that It includes program code. When the storage medium is run on an electronic device, the program code is used to enable the electronic device to execute the steps of the method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises computer instructions, which implement the steps of the method according to any one of claims 1 to 9 when the computer instructions are executed by a processor.

Citation Information

Patent Citations

  • Video classification method, equipment and medium

    CN112749685A

  • Systems and methods for partially supervised online action detection in untrimmed videos

    US20210357687A1