A method and apparatus for processing video

By obtaining the semantic features of the input statement and semantic enhancement of the video frame, the problem of difficulty in positioning video clips in the prior art is solved, and higher recognition accuracy and video content understanding ability are achieved.

CN113128285BActive Publication Date: 2025-06-17HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201911416325.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-31
Publication Date
2025-06-17
Estimated Expiration
2039-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to effectively locate video clips according to natural language descriptions, especially when processing complex video content.

Method used

By obtaining the semantic features of the input statement and semantic enhancement of the video frame based on these features, the semantic features are fused into the video features, thereby improving the accuracy of identifying the target video clips corresponding to the input statement.

Benefits of technology

It improves the accuracy of identifying the target video clips corresponding to the input statement, and enhances the understanding and matching ability of video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113128285B_ABST
    Figure CN113128285B_ABST
Patent Text Reader

Abstract

This application relates to the video clip localization technology in the field of computer vision in the field of artificial intelligence, and provides a method and device for processing videos. It relates to the field of artificial intelligence, specifically to the fields of computer vision and natural language processing. The method includes: obtaining the semantic features of the input statement; obtaining semantic enhancement for the video frames according to the semantic features to obtain the video features of the video frames, where the video features include the semantic features; determining whether the video clip to which the video frame belongs is the target video clip corresponding to the input statement according to the semantic features and the video features. This method helps to improve the accuracy of identifying the target video clip corresponding to the input statement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly, to a method and apparatus for processing video. Background Art

[0002] Artificial intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making. Research in the field of artificial intelligence includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, AI basic theory, etc.

[0003] With the rapid development of artificial intelligence technology, deep learning technology has made great progress in the fields of computer vision and natural language processing, and the joint research on the two fields has received increasing attention. For example, the problem of locating video clips according to natural language descriptions. However, compared with the problem of detecting static images according to natural language descriptions, the problem of locating video clips according to natural language descriptions is more complex.

[0004] Therefore, how to locate video clips according to natural language descriptions has become a technical problem that urgently needs to be solved. Summary of the Invention

[0005] This application provides a method and apparatus for processing video, which helps to improve the accuracy of identifying the target video clip corresponding to the input sentence.

[0006] In a first aspect, a method for processing video is provided. The method includes: obtaining the semantic feature of an input sentence; performing semantic enhancement on video frames according to the semantic feature to obtain video features of the video frames, where the video features include the semantic feature; and determining whether the video clip to which the video frame belongs is the target video clip corresponding to the input sentence according to the semantic feature and the video features.

[0007] In the embodiments of the present application, semantic enhancement is performed on a video frame according to the semantic feature to obtain the video feature of the video frame. The semantics corresponding to the input statement can be incorporated into the video feature of the video frame. At this time, according to the semantic feature and the video feature, the target video segment corresponding to the input statement is recognized, which can improve the accuracy of recognizing the target video segment corresponding to the input statement.

[0008] Among them, the semantic feature of the input statement can be the feature vector of the input statement, and the feature vector of the input statement can be used to represent the input statement. In other words, the semantic feature of the input statement can also be regarded as the vector form expression of the input statement.

[0009] For example, a recurrent neural network (RNN) can be used to obtain the semantic feature of the input statement. Alternatively, other neural networks can also be used to obtain the semantic feature of the input statement, which is not limited in the embodiments of the present application.

[0010] Similarly, the video feature of the video frame can be the feature vector of the video frame, and the feature vector of the video frame can be used to represent the video frame. In other words, the video feature of the video frame can also be regarded as the vector form expression of the video frame.

[0011] Among them, the fact that the semantic feature is included in the video feature may mean that the semantics corresponding to the input statement are included in the video feature, or the semantics corresponding to the input statement are carried in the video feature.

[0012] It should be noted that the above semantic enhancement may refer to collaboratively constructing the video feature of the video frame based on the semantic feature, or in other words, fusing the semantic feature (or it can also be understood as the semantics corresponding to the input statement) into the video feature of the video frame.

[0013] For example, when extracting the video feature of the video frame, semantic enhancement can be performed on the video frame based on the semantic feature to directly obtain the semantically enhanced (video feature of the) video frame.

[0014] For another example, the initial video feature of the video frame can be obtained first, and then semantic enhancement is performed on the initial video feature of the video frame based on the semantic feature to obtain the semantically enhanced (video feature of the) video frame.

[0015] In combination with the first aspect, in some implementations of the first aspect, the semantic enhancement of the video frame according to the semantic feature to obtain the video feature of the video frame includes: determining the word corresponding to the video frame in the input statement; and performing semantic enhancement on the video frame according to the semantic feature of the word corresponding to the video frame to obtain the video feature of the video frame.

[0016] In the embodiments of the present application, using the semantic feature of the word most relevant to the video frame in the input statement to perform semantic enhancement on the video frame can make the video feature of the video frame more accurate. At this time, identifying the target video segment corresponding to the input statement according to the video feature can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0017] In combination with the first aspect, in some implementations of the first aspect, the semantic enhancement of the video frame according to the semantic feature to obtain the video feature of the video frame in the video includes: extracting the feature of the video frame according to the semantic feature to obtain the video feature of the video frame.

[0018] In the embodiments of the present application, combining the semantic feature of the input statement to extract the feature of the video frame can directly perform semantic enhancement on the video feature of the video frame during the feature extraction process, which helps to improve the efficiency of identifying the target video segment corresponding to the input statement.

[0019] In combination with the first aspect, in some implementations of the first aspect, the method further includes: obtaining the initial video feature of the video frame; wherein, the semantic enhancement of the video frame according to the semantic feature to obtain the video feature of the video frame includes: performing semantic enhancement on the initial video feature according to the semantic feature to obtain the video feature of the video frame.

[0020] In combination with the first aspect, in some implementations of the first aspect, the method further includes: using the video features of at least one other video frame to perform feature fusion on the video feature of the video frame to obtain the fused video feature of the video frame, where the other video frame and the video frame belong to the same video; wherein, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the video feature includes: determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the fused video feature.

[0021] In an embodiment of the present application, by using the video features of other video frames in the video to perform feature fusion on the video features of the video frame and integrating the context information in the video into the video features of the video frame, the video features of the video frame can be made more accurate. At this time, identifying the target video segment corresponding to the input statement according to the video features can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0022] Optionally, the video features of the at least one other video frame can be added to the video features of the video frame, and the resulting sum is the fused video features of the video frame.

[0023] At this time, it can be considered that the video features of the at least one other video frame are fused in the fused video features of the video frame.

[0024] In an embodiment of the present application, the video features of all other video frames in the video except the video frame can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0025] Alternatively, the video features of all video frames (including the video frame) in the video can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0026] For example, the average value of the video features of all video frames (including the video frame) in the video can be calculated, and this average value is added to the video features of the video frame, and the resulting sum is the fused video features of the video frame.

[0027] For another example, the video includes a total of t video frames, and the video features of these t video frames can form the video feature sequence {f1, f2,..., f t}, where f j represents the semantic feature of the j-th word in the input statement, j is a positive integer less than t, and t is a positive integer. Multiply the video features in the video feature sequence {f1, f2,..., f t} pairwise (matrix multiplication) to obtain matrix B (matrix B can be called the correlation matrix, and the elements in matrix B can be called correlation features). Select a correlation feature for the video frame f j in matrix B, and add this correlation feature to the video features of the video frame f j in the video, and the resulting sum is the fused video features of the video frame f j in the video.

[0028] In combination with the first aspect, in some implementations of the first aspect, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the video feature includes: determining the hierarchical structure of the video segment in the time domain based on the video feature; and determining whether the video segment is the target video segment corresponding to the input statement according to the semantic feature and the hierarchical structure.

[0029] In the embodiments of the present application, using the video feature to determine the hierarchical structure of the video segment in the time domain can expand the receptive field of each video frame in the video segment while maintaining the size of the video feature of each video frame. At this time, identifying the target video segment corresponding to the input statement according to the semantic feature and the hierarchical structure can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0030] Optionally, one-dimensional dilated convolution or one-dimensional convolution can be used to determine the hierarchical structure of the video segment in the time domain based on the video feature.

[0031] In a second aspect, a device for processing video is provided, including: obtaining the semantic feature of an input statement; semantically enhancing a video frame according to the semantic feature to obtain the video feature of the video frame, where the video feature includes the semantic feature; and determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the video feature.

[0032] In the embodiments of the present application, semantically enhancing the video frame according to the semantic feature to obtain the video feature of the video frame can incorporate the semantics corresponding to the input statement into the video feature of the video frame. At this time, identifying the target video segment corresponding to the input statement according to the semantic feature and the video feature can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0033] Wherein, the semantic feature of the input statement may be the feature vector of the input statement, and the feature vector of the input statement can be used to represent the input statement. In other words, the semantic feature of the input statement can also be considered as the vector form expression of the input statement.

[0034] For example, a recurrent neural network (RNN) can be used to obtain the semantic feature of the input statement. Alternatively, other neural networks can also be used to obtain the semantic feature of the input statement, and the embodiments of the present application do not limit this.

[0035] Similarly, the video feature of the video frame can be the feature vector of the video frame, and the feature vector of the video frame can be used to represent the video frame. In other words, the video feature of the video frame can also be regarded as the vector form expression of the video frame.

[0036] Among them, that the semantic feature is included in the video feature may mean that the semantics corresponding to the input statement is included in the video feature, or the semantics corresponding to the input statement is carried in the video feature.

[0037] It should be noted that the above semantic enhancement may refer to co-constructing the video feature of the video frame based on the semantic feature, or in other words, fusing the semantic feature (which can also be understood as the semantics corresponding to the input statement) into the video feature of the video frame.

[0038] For example, when extracting the video feature of the video frame, semantic enhancement can be performed on the video frame based on the semantic feature to directly obtain the semantically enhanced video feature (of the video frame).

[0039] For another example, the initial video feature of the video frame can also be obtained first, and then semantic enhancement is performed on the initial video feature of the video frame based on the semantic feature to obtain the semantically enhanced video feature (of the video frame).

[0040] Combined with the second aspect, in some implementation manners of the second aspect, the semantic enhancement of the video frame according to the semantic feature to obtain the video feature of the video frame includes: determining the word corresponding to the video frame in the input statement; performing semantic enhancement on the video frame according to the semantic feature of the word corresponding to the video frame to obtain the video feature of the video frame.

[0041] In the embodiments of the present application, using the semantic feature of the word most relevant to the video frame in the input statement to perform semantic enhancement on the video frame can make the video feature of the video frame more accurate. At this time, identifying the target video segment corresponding to the input statement according to the video feature can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0042] Combined with the second aspect, in some implementation manners of the second aspect, the semantic enhancement of the video frame according to the semantic feature to obtain the video feature of the video frame in the video includes: performing feature extraction on the video frame according to the semantic feature to obtain the video feature of the video frame.

[0043] In the embodiments of the present application, by extracting features from the video frames in combination with the semantic features of the input statement, semantic enhancement can be directly performed on the video features of the video frames during the feature extraction process, which helps to improve the efficiency of identifying the target video segment corresponding to the input statement.

[0044] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: obtaining the initial video features of the video frames; wherein, the semantic enhancement of the video frames according to the semantic features to obtain the video features of the video frames includes: performing semantic enhancement on the initial video features according to the semantic features to obtain the video features of the video frames.

[0045] In combination with the second aspect, in some implementation manners of the second aspect, the method further includes: using the video features of at least one other video frame to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame, where the other video frame and the video frame belong to the same video; wherein, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features includes: determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the fused video features.

[0046] In the embodiments of the present application, using the video features of other video frames in the video to perform feature fusion on the video features of the video frame and integrating the context information in the video into the video features of the video frame can make the video features of the video frame more accurate. At this time, identifying the target video segment corresponding to the input statement according to the video features can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0047] Optionally, the video features of the at least one other video frame can be added to the video features of the video frame, and the result after addition is the fused video features of the video frame.

[0048] At this time, it can be considered that the video features of the at least one other video frame are fused in the fused video features of the video frame.

[0049] In the embodiments of the present application, the video features of all other video frames in the video except the video frame can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0050] Alternatively, the video features of all video frames (including the video frame) in the video can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0051] For example, the average value of the video features of all video frames (including the video frame) in the video can be calculated, and the average value is added to the video features of the video frame. The result after addition is the fused video feature of the video frame.

[0052] For another example, the video contains t video frames in total. The video features of these t video frames can form the video feature sequence {f1, f2,..., ft} of the video, where fj represents the semantic feature of the j-th word in the input statement, j is a positive integer less than t, and t is a positive integer. The video features in the video feature sequence {f1, f2,..., ft} are multiplied pairwise (matrix multiplication) to obtain matrix B (matrix B can be called the correlation matrix, and the elements in matrix B can be called correlation features). A correlation feature is selected for the video frame fj in the video from matrix B, and the correlation feature is added to the video features of the video frame fj in the video. The result after addition is the fused video feature of the video frame fj in the video.

[0053] In combination with the second aspect, in some implementation manners of the second aspect, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the video feature includes: determining the hierarchical structure of the video segment in the time domain based on the video feature; and determining whether the video segment is the target video segment corresponding to the input statement according to the semantic feature and the hierarchical structure.

[0054] In the embodiments of the present application, using the video feature to determine the hierarchical structure of the video segment in the time domain can expand the receptive field of each video frame in the video segment while maintaining the size of the video features of each video frame. At this time, identifying the target video segment corresponding to the input statement according to the semantic feature and the hierarchical structure can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0055] Optionally, one-dimensional dilated convolution or one-dimensional convolution can be used to determine the hierarchical structure of the video segment in the time domain based on the video feature.

[0056] In a third aspect, a device for processing video is provided. The device includes: a memory for storing a program; and a processor for executing the program stored in the memory. When the program stored in the memory is executed, the processor is used to execute the method in any one of the implementation manners in the first aspect above.

[0057] The processor in the third aspect above can be either a central processing unit (CPU) or a combination of a CPU and a neural network computing processor. The neural network computing processor here can include a graphics processing unit (GPU), a neural-network processing unit (NPU), a tensor processing unit (TPU), and so on. Among them, the TPU is an application-specific integrated circuit of artificial intelligence accelerator fully customized by Google for machine learning.

[0058] In a fourth aspect, a computer-readable medium is provided. The computer-readable medium stores program code for a device to execute, and the program code includes a method for executing any implementation in the first aspect.

[0059] In a fifth aspect, a computer program product containing instructions is provided. When the computer program product runs on a computer, it causes the computer to execute the method in any implementation in the first aspect above.

[0060] In a sixth aspect, a chip is provided. The chip includes a processor and a data interface. The processor reads instructions stored on a memory through the data interface and executes the method in any implementation in the first aspect above.

[0061] Optionally, as an implementation, the chip may further include a memory. Instructions are stored in the memory, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to execute the method in any implementation in the first aspect.

[0062] The above chip may specifically be a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC).

[0063] In a seventh aspect, an electronic device is provided. The electronic device includes a device for processing video in any aspect in the second aspect above.

[0064] When the above electronic device includes a device for processing video in any aspect in the second aspect above, the electronic device may specifically be a terminal device or a server.

[0065] In the embodiments of the present application, semantic enhancement is performed on video frames according to the semantic features to obtain the video features of the video frames. The semantics corresponding to the input statement can be incorporated into the video features of the video frames. At this time, according to the semantic features and the video features, the target video segment corresponding to the input statement is recognized, which can improve the accuracy of recognizing the target video segment corresponding to the input statement. Description of the Drawings

[0066] Figure 1 It is a schematic diagram of an artificial intelligence main framework provided by an embodiment of the present application.

[0067] Figure 2 It is a schematic structural diagram of a system architecture provided by an embodiment of the present application.

[0068] Figure 3 It is a schematic structural diagram of a convolutional neural network provided by an embodiment of the present application.

[0069] Figure 4 It is a schematic structural diagram of another convolutional neural network provided by an embodiment of the present application.

[0070] Figure 5 It is a schematic hardware structure diagram of a chip provided by an embodiment of the present application.

[0071] Figure 6 It is a schematic diagram of a system architecture provided by an embodiment of the present application.

[0072] Figure 7 It is a schematic flowchart of a method for processing video according to an embodiment of the present application.

[0073] Figure 8 It is a schematic flowchart of a method for processing video according to another embodiment of the present application.

[0074] Figure 9 It is a schematic hardware structure diagram of a device for processing video according to an embodiment of the present application.

[0075] Figure 10 It is a schematic hardware structure diagram of a device for processing video according to an embodiment of the present application. Detailed Embodiments

[0076] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0077] Figure 1Shows a schematic diagram of an artificial intelligence agent framework, which describes the overall workflow of an artificial intelligence system and is applicable to the general requirements of the artificial intelligence field.

[0078] The above artificial intelligence theme framework is elaborated in detail from two dimensions: the "intelligent information chain" (horizontal axis) and the "information technology (IT) value chain" (vertical axis).

[0079] The "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes the refinement process of "data - information - knowledge - wisdom".

[0080] The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.

[0081] (1) Infrastructure:

[0082] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported through the basic platform.

[0083] The infrastructure can communicate with the external through sensors, and the computing power of the infrastructure can be provided by intelligent chips.

[0084] The intelligent chips here can be hardware acceleration chips such as central processing unit (CPU), neural-network processing unit (NPU), graphics processing unit (GPU), application specific integrated circuit (ASIC), and field programmable gate array (FPGA).

[0085] The basic platform of the infrastructure can include relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc.

[0086] For example, for the infrastructure, data can be obtained through communication with the external through sensors, and then these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for calculation.

[0087] (2) Data:

[0088] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. This data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0089] (3) Data processing:

[0090] The above data processing usually includes data training, machine learning, deep learning, search, inference, decision-making and other processing methods.

[0091] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0092] Inference refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information for machine thinking and problem-solving according to the inference control strategy. The typical function is search and matching.

[0093] Decision-making refers to the process of making decisions after the intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.

[0094] (4) General capabilities:

[0095] After the data is processed through the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0096] (5) Intelligent products and industry applications:

[0097] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which are the encapsulation of the overall artificial intelligence solution, productize the intelligent information decision-making, and realize the landing application. Its application fields mainly include: intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city, intelligent terminal, etc.

[0098] The embodiments of this application can be applied in many fields of artificial intelligence. For example, intelligent manufacturing, intelligent transportation, smart home, intelligent healthcare, intelligent security, autonomous driving, safe city and other fields.

[0099] Specifically, the embodiments of this application can be specifically applied to the management and retrieval of the multimedia library in the cloud, or can also be applied to the management and retrieval of the multimedia library at the terminal, or can also be applied to other scenarios where natural language is used to manage and retrieve the multimedia library containing a large number of videos.

[0100] The following briefly introduces the application scenario of using natural language to search for interesting video clips (in a multimedia library).

[0101] Video clip search:

[0102] The multimedia library contains a large number of videos. When a user searches for interesting video clips, using natural language for query can improve the interaction method (for video clip search), making the management and retrieval of videos by the user more convenient and enhancing the user experience.

[0103] Specifically, when a user stores a large number of videos on a terminal device (such as a mobile phone) or in the cloud, by inputting natural language, the user can query interesting video clips, or can classify and manage the stored videos, thereby improving the user experience.

[0104] For example, by adopting the method for processing videos in the embodiments of the present application, a device (or model) suitable for processing videos can be constructed. When a user hopes to search for video clips about "a baby is eating" in the multimedia library, the natural sentence "a baby is eating" can be input. By inputting the videos in the multimedia library and the input natural sentence into the above-constructed device (or model), the video clips about "a baby is eating" in the multimedia library can be obtained, thus completing the search for interesting video clips.

[0105] Since the embodiments of the present application involve a large number of applications of neural networks, for the convenience of understanding, the following first introduces the relevant terms and concepts of neural networks that may be involved in the embodiments of the present application.

[0106] (1) Neural network

[0107] A neural network can be composed of neural units. A neural unit can refer to an operation unit with x s and intercept 1 as inputs, and the output of this operation unit can be:

[0108]

[0109] where s = 1, 2,..., n, and n is a natural number greater than 1, and W s is x sThe weight of, b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non - linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting multiple such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, and the local receptive field can be a region composed of several neurons.

[0110] (2) Deep neural network

[0111] A deep neural network (DNN), also known as a multi - layer neural network, can be understood as a neural network with multiple hidden layers. Classifying the DNN according to the positions of different layers, the neural network inside the DNN can be divided into three categories: the input layer, the hidden layer, and the output layer. Generally, the first layer is the input layer, the last layer is the output layer, and the middle layers are all hidden layers. The layers are fully connected, that is, any neuron in the i - th layer must be connected to any neuron in the (i + 1) - th layer.

[0112] Although the DNN looks very complex, in terms of the work of each layer, it is actually not complex. Simply put, it is the following linear relationship expression: Among them, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also known as the coefficient), and α() is the activation function. Each layer simply performs such a simple operation on the input vector to obtain the output vector Due to the large number of layers in the DNN, the number of coefficients W and the offset vector is also relatively large. The definitions of these parameters in the DNN are as follows: Taking the coefficient W as an example: Suppose in a three - layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer where the coefficient W is located, and the subscripts correspond to the index 2 of the output third layer and the index 4 of the input second layer.

[0113] In summary, the coefficient from the k - th neuron in the (L - 1) - th layer to the j - th neuron in the L - th layer is defined as

[0114] It should be noted that there is no W parameter in the input layer. In a deep neural network, more hidden layers enable the network to better depict complex situations in the real world. In theory, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is also a process of learning the weight matrix, and its ultimate goal is to obtain the weight matrices of all layers of the trained deep neural network (the weight matrix formed by vectors W of many layers).

[0115] (3) Convolutional Neural Network

[0116] A convolutional neural network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor composed of convolutional layers and subsampling layers, and this feature extractor can be regarded as a filter. A convolutional layer refers to the neuron layer in a convolutional neural network that performs convolutional processing on the input signal. In the convolutional layer of a convolutional neural network, a neuron can be connected to only some neighboring layer neurons. In a convolutional layer, there are usually several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weight here is the convolutional kernel. Sharing weights can be understood as a way of extracting image information that is independent of position. The convolutional kernel can be initialized in the form of a matrix of random size, and during the training process of the convolutional neural network, the convolutional kernel can learn to obtain reasonable weights. Additionally, the direct benefit brought by sharing weights is to reduce the connections between layers of the convolutional neural network while also reducing the risk of overfitting.

[0117] (4) Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, it is from the input layer to the hidden layer and then to the output layer, and there are full connections between layers, while there are no connections between individual nodes within each layer. Although this ordinary neural network has solved many problems, it is still powerless in many aspects. For example, when you want to predict what the next word in a sentence is, you generally need to use the previous words because the words before and after in a sentence are not independent. The reason why RNN is called a recurrent neural network is that the current output of a sequence is also related to the previous output. The specific manifestation is that the network will remember the previous information and apply it to the calculation of the current output, that is, the nodes within the hidden layer itself are no longer unconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNN can process sequential data of any length. The training of RNN is the same as that of traditional CNN or DNN.

[0118] Since we already have convolutional neural networks, why do we still need recurrent neural networks? The reason is simple. In convolutional neural networks, there is a premise assumption that elements are independent of each other, and the input and output are also independent, such as cats and dogs. However, in the real world, many elements are interconnected. For example, the change of stocks over time, or a person says: "I like traveling, and my favorite place is Yunnan. I must go there if I have the chance in the future." Here, fill in the blank, and humans should all know that it is "Yunnan". Because humans can make inferences based on the context. But how can we make machines do this? RNN came into being. RNN aims to enable machines to have the ability to remember like humans. Therefore, the output of RNN needs to depend on the current input information and historical memory information.

[0119] (5) Loss function

[0120] During the process of training a deep neural network, since we hope that the output of the deep neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two (of course, there is usually an initialization process before the first update, that is, pre-configure parameters for each layer in the deep neural network). For example, if the predicted value of the network is too high, we adjust the weight vector to make it predict lower, and keep adjusting until the deep neural network can predict the real target value or a value very close to the real target value. Therefore, we need to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.

[0121] (6) Backpropagation algorithm

[0122] Neural networks can use the error backpropagation (BP) algorithm to correct the size of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the error loss information is propagated backward to update the parameters in the initial neural network model, so that the error loss converges. The backpropagation algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.

[0123] As Figure 2 shown, an embodiment of the present application provides a system architecture 100. InFigure 2 Among them, the data acquisition device 160 is used to acquire training data. For the method of processing videos according to the embodiments of the present application, the training data may include input statements, training videos, and video segments in the training videos that have the highest matching degree with the input statements. Among them, the video segments in the training videos that have the highest matching degree with the input statements may be pre-annotated video segments manually.

[0124] After the training data is acquired, the data acquisition device 160 stores the training data in the database 130, and the training device 120 trains to obtain the target model / rule 101 based on the training data maintained in the database 130.

[0125] The following describes how the training device 120 obtains the target model / rule 101 based on the training data. The training device 120 processes the training videos based on the input statements, and compares the output video segments with the video segments in the training videos that have the highest matching degree with the input statements until the difference between the video output by the training device 120 and the video segments in the training videos that have the highest matching degree with the input statements is less than a certain threshold, thereby completing the training of the target model / rule 101.

[0126] The above target model / rule 101 can be used to implement the method of processing videos according to the embodiments of the present application. The target model / rule 101 in the embodiments of the present application may specifically be a device (or model) for processing videos in the embodiments of the present application, and the device (or model) for processing videos may include multiple neural networks. It should be noted that in actual applications, the training data maintained in the database 130 does not necessarily come from the acquisition of the data acquisition device 160, and it may also be received from other devices. Additionally, it should be noted that the training device 120 does not necessarily train the target model / rule 101 entirely based on the training data maintained in the database 130, and it may also obtain training data from the cloud or other places for model training. The above description should not be regarded as a limitation to the embodiments of the present application.

[0127] The target model / rule 101 trained according to the training device 120 can be applied to different systems or devices, such as applied to Figure 2 the execution device 110 shown in the figure. The execution device 110 may be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, an augmented reality (AR) / virtual reality (VR) device, a vehicle-mounted terminal, etc., or it may also be a server or the cloud, etc. In Figure 2In this case, the execution device 110 configures an input / output (I / O) interface 112 for data interaction with external devices. A user can input data to the I / O interface 112 through a client device 140. The input data in the embodiments of this application may include: videos and input statements input by the client device.

[0128] The preprocessing modules 113 and 114 are used to preprocess the input data (such as the input videos and input statements) received by the I / O interface 112. In the embodiments of this application, the preprocessing modules 113 and 114 may not exist (or only one of the preprocessing modules may exist), and the computing module 111 may directly process the input data.

[0129] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing, the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.

[0130] Finally, the I / O interface 112 returns the processing result, such as the obtained video clip, to the client device 140 for the user.

[0131] It should be noted that the training device 120 can generate corresponding target models / rules 101 based on different training data for different targets or tasks. The corresponding target models / rules 101 can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user.

[0132] In Figure 2 the shown case, the user can manually provide input data, and this manual provision can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is required, the user can set corresponding permissions in the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific presentation forms can be display, sound, action and other specific ways. The client device 140 can also be used as a data acquisition end to collect the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as new sample data and store them in the database 130. Of course, the collection can also be directly performed by the I / O interface 112 without going through the client device 140, and the input data input to the I / O interface 112 and the output results of the output I / O interface 112 as shown are stored in the database 130 as new sample data.

[0133] It should be noted that Figure 2 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationships among the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in Figure 2 , the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110.

[0134] As Figure 2 shown, the target model / rule 101 is trained according to the training device 120. The target model / rule 101 can be a device (or model) for processing videos in the present application in the embodiments of the present application. The device (or model) for processing videos can include multiple neural networks. Specifically, the device (or model) for processing videos can include a CNN, deep convolutional neural networks (DCNN), recurrent neural network (RNN), and so on.

[0135] Since the CNN is a very common neural network, the structure of the CNN will be introduced in detail below in combination with Figure 3 . As described in the basic concept introduction above, the convolutional neural network is a deep neural network with a convolutional structure and is a deep learning architecture. The deep learning architecture refers to performing multiple levels of learning at different abstraction levels through machine learning algorithms. As a deep learning architecture, the CNN is a feed-forward artificial neural network, and each neuron in the feed-forward artificial neural network can respond to the input image.

[0136] The structure of the convolutional neural network specifically adopted in the embodiments of the present application can be as Figure 3 shown. In Figure 3 , the convolutional neural network (CNN) 200 can include an input layer 210, a convolutional layer / pooling layer 220 (where the pooling layer is optional), and a neural network layer 230.

[0137] In the embodiments of the present application, a video frame can be considered as an image. Therefore, taking the processing of an image as an example, the structure of the convolutional neural network will be introduced. For example, the input layer 210 can obtain the image to be processed and hand over the obtained image to be processed to the convolutional layer / pooling layer 220 and the subsequent neural network layer 230 for processing, and the processing result of the image can be obtained. The internal layer structure of the CNN 200 in Figure 3 will be introduced in detail below.

[0138] Convolutional layer / pooling layer 220:

[0139] Convolutional layer:

[0140] As Figure 3 shown, the convolutional layer / pooling layer 220 may include layers such as examples 221 - 226. For example: in one implementation, layer 221 is a convolutional layer, layer 222 is a pooling layer, layer 223 is a convolutional layer, layer 224 is a pooling layer, 225 is a convolutional layer, and 226 is a pooling layer; in another implementation, 221 and 222 are convolutional layers, 223 is a pooling layer, 224 and 225 are convolutional layers, and 226 is a pooling layer. That is, the output of the convolutional layer can be used as the input of the subsequent pooling layer or as the input of another convolutional layer to continue the convolution operation.

[0141] Next, taking the convolutional layer 221 as an example, the internal working principle of one convolutional layer will be introduced.

[0142] The convolutional layer 221 may include many convolutional operators, which are also called kernels. Their role in image processing is equivalent to a filter that extracts specific information from the input image matrix. Essentially, a convolutional operator can be a weight matrix, which is usually predefined. During the convolution operation on the image, the weight matrix usually processes the input image pixel by pixel (or two pixels by two pixels... depending on the value of the stride) along the horizontal direction, so as to complete the work of extracting specific features from the image. The size of this weight matrix should be related to the size of the image. It should be noted that the depth dimension of the weight matrix and the depth dimension of the input image are the same. During the convolution operation, the weight matrix will extend to the entire depth of the input image. Therefore, convolving with a single weight matrix will produce a convolved output with a single depth dimension. However, in most cases, a single weight matrix is not used, but multiple weight matrices with the same size (row × column) are applied, that is, multiple matrices of the same type. The output of each weight matrix is stacked to form the depth dimension of the convolutional image, and here the dimension can be understood as determined by the above-mentioned "multiple". Different weight matrices can be used to extract different features from the image. For example, one weight matrix is used to extract image edge information, another weight matrix is used to extract specific colors of the image, and another weight matrix is used to blur the unwanted noise in the image, etc. The sizes (row × column) of these multiple weight matrices are the same, and the sizes of the convolutional feature maps extracted by these multiple weight matrices with the same size are also the same. Then, the multiple convolutional feature maps with the same size that are extracted are combined to form the output of the convolution operation.

[0143] The weight values ​​in these weight matrices need to be obtained through a lot of training in practical applications. The weight matrices formed by the weight values ​​obtained through training can be used to extract information from the input image, so that the convolutional neural network 200 can make correct predictions.

[0144] When the convolutional neural network 200 has multiple convolutional layers, the initial convolutional layer (for example, 221) often extracts more general features, which can also be called low-level features. As the depth of the convolutional neural network 200 increases, the features extracted by the later convolutional layers (for example, 226) become more and more complex, such as high-level semantic features. Features with higher semantics are more suitable for the problem to be solved.

[0145] Pooling layer:

[0146] Since it is often necessary to reduce the number of training parameters, it is often necessary to periodically introduce a pooling layer after the convolution layer. Figure 3 Each layer 221-226 illustrated in 220 may be a convolution layer followed by a pooling layer, or may be multiple convolution layers followed by one or more pooling layers. In the image processing process, the only purpose of the pooling layer is to reduce the spatial size of the image. The pooling layer may include an average pooling operator and / or a maximum pooling operator for sampling the input image to obtain an image of smaller size. The average pooling operator may calculate the pixel values ​​in the image within a specific range to generate an average value as the result of average pooling. The maximum pooling operator may take the pixel with the largest value in the range within a specific range as the result of maximum pooling. In addition, just as the size of the weight matrix used in the convolution layer should be related to the image size, the operator in the pooling layer should also be related to the image size. The size of the image output after processing by the pooling layer may be smaller than the size of the image input to the pooling layer, and each pixel in the image output by the pooling layer represents the average value or maximum value of the corresponding sub-region of the image input to the pooling layer.

[0147] Neural Network Layer 230:

[0148] After being processed by the convolution layer / pooling layer 220, the convolution neural network 200 is not sufficient to output the required output information. As mentioned above, the convolution layer / pooling layer 220 only extracts features and reduces the parameters brought by the input image. However, in order to generate the final output information (the required class information or other related information), the convolution neural network 200 needs to use the neural network layer 230 to generate one or a group of outputs of the required number of classes. Therefore, the neural network layer 230 may include multiple hidden layers (such as Figure 3231, 232 to 23n) as shown, and an output layer 240. The parameters included in the multi-layer hidden layer can be pre-trained according to the relevant training data of the specific task type. For example, the task type can include image recognition, image classification, image super-resolution reconstruction, and so on.

[0149] After the multi-layer hidden layer in the neural network layer 230, that is, the last layer of the entire convolutional neural network 200 is the output layer 240. The output layer 240 has a loss function similar to categorical cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network 200 (such as Figure 3 The propagation from 210 to 240 is the forward propagation) is completed, the backpropagation (such as Figure 3 The propagation from 240 to 210 is the backpropagation) will start to update the weight values and biases of the previously mentioned layers to reduce the loss of the convolutional neural network 200, that is, the error between the result output by the convolutional neural network 200 through the output layer and the ideal result.

[0150] The structure of the convolutional neural network adopted in the embodiments of the present application can be as Figure 4 shown. In Figure 4 , the convolutional neural network (CNN) 200 can include an input layer 110, a convolutional layer / pooling layer 120 (where the pooling layer is optional), and a neural network layer 130. Compared with Figure 3 , Figure 4 In the convolutional layer / pooling layer 120 in , multiple convolutional layers / pooling layers are parallel, and the features extracted separately are all input to the fully neural network layer 130 for processing.

[0151] It should be noted that Figure 3 and Figure 4 The convolutional neural networks shown are only examples of two possible convolutional neural networks adopted in the embodiments of the present application. In specific applications, the convolutional neural network adopted in the embodiments of the present application can also exist in the form of other network models.

[0152] Figure 5 This is the hardware structure of a chip provided by the embodiments of the present application. The chip includes a neural network processor 50. The chip can be set in an execution device 110 as shown in Figure 1 to complete the computing work of the computing module 111. The chip can also be set in a training device 120 as shown in Figure 1 to complete the training work of the training device 120 and output the target model / rule 101. The algorithms of each layer in the convolutional neural network as shown in Figure 3 and Figure 4 can all be implemented in the chip as shown in Figure 5 .

[0153] The neural network processor NPU 50 is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are assigned by the main CPU. The core part of the NPU is the arithmetic circuit 503, and the controller 504 controls the arithmetic circuit 503 to extract data from the memory (weight memory or input memory) and perform operations.

[0154] In some implementations, the arithmetic circuit 503 includes multiple processing engines (PEs) inside. In some implementations, the arithmetic circuit 503 is a two-dimensional systolic array. The arithmetic circuit 503 can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 503 is a general matrix processor.

[0155] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 502 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 501 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator 508.

[0156] The vector calculation unit 507 can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector calculation unit 507 can be used for network calculations in non-convolutional / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.

[0157] In some implementations, the vector calculation unit 507 can store the processed output vector in the unified buffer 506. For example, the vector calculation unit 507 can apply a non-linear function to the output of the arithmetic circuit 503, such as a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 507 generates normalized values, combined values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 503, for example, for use in subsequent layers in a neural network.

[0158] The unified memory 506 is used to store input data and output data.

[0159] The weight data directly transfers the input data in the external memory to the input memory 501 and / or the unified memory 506 through the storage unit access controller 505 (direct memory access controller, DMAC), stores the weight data in the external memory into the weight memory 502, and stores the data in the unified memory 506 into the external memory.

[0160] The bus interface unit (BIU) 510 is used to interact between the main CPU, DMAC, and the instruction fetch memory 509 through the bus.

[0161] The instruction fetch buffer 509 connected to the controller 504 is used to store the instructions used by the controller 504.

[0162] The controller 504 is used to call the instructions cached in the instruction memory 509 to control the working process of the arithmetic accelerator.

[0163] Generally, the unified memory 506, the input memory 501, the weight memory 502, and the instruction fetch memory 509 are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDR SDRAM for short), a high bandwidth memory (HBM), or other readable and writable memories.

[0164] Among them, Figure 3 and Figure 4 The operations of each layer in the convolutional neural network shown can be executed by the arithmetic circuit 503 or the vector calculation unit 507.

[0165] The execution device 110 introduced above Figure 2 in can execute each step of the method for processing video in the embodiments of the present application. Figure 3 and Figure 4 The CNN model shown and Figure 5 The chip shown can also be used to execute each step of the method for processing video in the embodiments of the present application. The method for processing video in the embodiments of the present application will be introduced in detail below with reference to the accompanying drawings.

[0166] As Figure 6As shown in the figure, an embodiment of the present application provides a system architecture 300. The system architecture includes local devices 301, local device 302, an execution device 210, and a data storage system 250. Among them, the local device 301 and the local device 302 are connected to the execution device 210 through a communication network.

[0167] The execution device 210 can be implemented by one or more servers. Optionally, the execution device 210 can be used in cooperation with other computing devices, such as devices like data memories, routers, load balancers, etc. The execution device 210 can be arranged on one physical site or distributed across multiple physical sites. The execution device 210 can use the data in the data storage system 250 or call the program code in the data storage system 250 to implement the method for processing videos in the embodiments of the present application.

[0168] Specifically, the execution device 210 can perform the following processes: obtain the semantic features of the input statement; semantically enhance the video frames according to the semantic features to obtain the video features of the video frames, where the video features include the semantic features; and determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features.

[0169] Through the above processes, the execution device 210 can be built into a device (or model) for processing videos. The device (or model) for processing videos can include one or more neural networks, and the device (or model) for processing videos can be used for video segment search or positioning, retrieval or management of multimedia libraries, and so on.

[0170] Users can operate their respective user devices (such as local device 301 and local device 302) to interact with the execution device 210. Each local device can represent any computing device, such as a personal computer, a computer workstation, a smart phone, a tablet computer, a smart camera, a smart car, or other types of cellular phones, media consumption devices, wearable devices, set-top boxes, game consoles, etc.

[0171] The local device of each user can interact with the execution device 210 through a communication network using any communication mechanism / communication standard. The communication network can be a wide area network, a local area network, a point-to-point connection, or any combination thereof.

[0172] In one implementation, the local device 301 and the local device 302 obtain the relevant parameters of the device (or model) for processing videos from the execution device 210, deploy the device (or model) for processing videos on the local device 301 and the local device 302, and use the device (or model) for processing videos for video segment search or positioning, retrieval or management of multimedia libraries, and so on.

[0173] In another implementation, a device (or model) for processing video can be directly deployed on the execution device 210. The execution device 210 obtains the input video and input statements from the local device 301 and the local device 302, and performs operations such as searching or locating video segments and retrieving or managing multimedia according to the device (or model) for processing video.

[0174] The above-mentioned execution device 210 can also be a cloud device. In this case, the execution device 210 can be deployed in the cloud; or, the above-mentioned execution device 210 can also be a terminal device. In this case, the execution device 210 can be deployed on the user terminal side. The embodiments of the present application do not limit this.

[0175] In the present application, a video is a sequence of video frames composed of video frames. Each video frame in it can also be regarded as a picture or an image. That is to say, a video can also be regarded as a sequence of pictures composed of pictures, or a sequence of images composed of images.

[0176] A video can also be divided into one or more video clips, and each video clip is composed of one or more video frames. For example, it can be divided into multiple video clips according to the content in the video, or it can also be divided into multiple video clips according to the time coordinates in the video. In the embodiments of the present application, the method of dividing a video into video clips is not limited.

[0177] The technical solutions of the embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0178] Figure 7 It is a schematic flowchart of a method for processing video according to the present application. The method 700 for processing video may include S710, S720, and S730. In some examples, the method for processing video may be executed by Figure 2 the execution device 110 in Figure 5 the chip shown in Figure 6 the execution device 210 in

[0179] S710, obtaining the semantic features of the input statement.

[0180] Optionally, a neural network can be used to obtain the semantic features (or semantic information) of the input statement.

[0181] Among them, the semantic features of the input statement can be the feature vector of the input statement, and the feature vector of the input statement can be used to represent the input statement. In other words, the semantic features of the input statement can also be regarded as the vector form expression of the input statement.

[0182] For example, a recurrent neural network (RNN) can be used to obtain the semantic features of an input statement. Alternatively, other neural networks can also be used to obtain the semantic features of the input statement, and the embodiments of the present application do not limit this.

[0183] Optionally, a neural network can be used to obtain the semantic features of each word (or term) in the input statement.

[0184] Correspondingly, the input statement can be represented as a semantic feature sequence composed of the semantic features of the respective words in the input statement.

[0185] For example, if the input statement includes k words (or terms), and a neural network is used to obtain the semantic features of the k words in the input statement, then the semantic feature sequence {w1, w2, …, w k} composed of the semantic features of the k words can be used to represent the semantic features of the input statement, where w i represents the semantic feature of the i-th word in the input statement, i is a positive integer less than k, and k is a positive integer.

[0186] S720, perform semantic enhancement on the video frame according to the semantic features to obtain the video features of the video frame.

[0187] Among them, the semantic features are included in the video features. Or in other words, the semantics corresponding to the input statement are included in the video features, or the semantics corresponding to the input statement are carried in the video features.

[0188] Optionally, in the embodiments of the present application, it is also possible to first determine the word in the input statement corresponding to the video frame; perform semantic enhancement on the video frame according to the semantic features of the word corresponding to the video frame to obtain the video features of the video frame.

[0189] The word in the input statement corresponding to the video frame mentioned here can be considered as: the word in the input statement that is most relevant to the content corresponding to the video frame. Or, it can also be considered that: among the multiple words included in the input statement, the semantics of this word are most relevant to the content corresponding to the video frame.

[0190] In the embodiments of the present application, using the semantic features of the word in the input statement that is most relevant to the video frame to perform semantic enhancement on the video frame can make the video features of the video frame more accurate. At this time, identifying the target video segment corresponding to the input statement according to the video features can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0191] The method for determining the word in the input statement corresponding to the video frame described above can be specifically as follows Figure 8As shown in the embodiments herein, details are not described again here.

[0192] It should be noted that the semantic enhancement in the above S720 may refer to jointly constructing the video features of the video frame based on the semantic features, or in other words, integrating the semantic features (or which can also be understood as the semantics corresponding to the input statement) into the video features of the video frame.

[0193] Optionally, in the embodiments of the present application, the semantic enhancement of the video frame according to the semantic features can be implemented in the following several ways:

[0194] Method 1:

[0195] In the embodiments of the present application, when extracting the video features of the video frame, the video frame can be semantically enhanced based on the semantic features to directly obtain the semantically enhanced (video frame's) video features.

[0196] Optionally, the semantic enhancement of the video frame according to the semantic features to obtain the video features of the video frame in the video may include: extracting the video features of the video frame according to the semantic features.

[0197] Generally, a pre-trained neural network is used to extract the video features of the video frame. According to the method in the embodiments of the present application, the semantic features can be integrated into the neural network, and when training the neural network, the video features of the video frame are extracted based on the semantic features.

[0198] After the neural network is trained, when extracting the video features of the video frame, the video frame can be semantically enhanced based on the semantic features, that is, the video features of the video frame are extracted according to the semantic features to obtain the video features of the video frame.

[0199] At this time, the semantic (or the semantics of the words in the input statement corresponding to the video frame) corresponding to the input statement is included in the video features obtained after feature extraction of the video frame, or in other words, the video features carry the semantic (or the semantics of the words in the input statement corresponding to the video frame) corresponding to the input statement.

[0200] In Method 1, by combining the semantic features of the input statement to extract the features of the video frame, the video features of the video frame can be directly semantically enhanced during the feature extraction process, which helps to improve the efficiency of identifying the target video segment corresponding to the input statement.

[0201] Method 2:

[0202] In an embodiment of the present application, the initial video features of the video frame may also be obtained first, and then the initial video features of the video frame are semantically enhanced based on the semantic features to obtain the video features of the video frame after semantic enhancement.

[0203] Optionally, before S720 above, the method 700 may further include S722.

[0204] S722, obtaining the initial video features of the video frame.

[0205] The initial video features of the video frame may be the feature vector of the video frame, and the feature vector of the video frame can be used to represent the video frame. In other words, the initial video features of the video frame can also be considered as the vector form expression of the video frame.

[0206] At this time, performing semantic enhancement on the video frame according to the semantic features to obtain the video features of the video frame may include: performing semantic enhancement on the initial video features according to the semantic features to obtain the video features of the video frame.

[0207] For example, a convolution kernel may be determined based on the semantic features, and the initial video features are convolved using the convolution kernel to fuse the semantics corresponding to the input statement into the video features of the video frame, thereby realizing semantic enhancement of the video frame.

[0208] The specific method of performing convolution processing on the initial video features based on the semantic features may be as shown in the embodiments below, which will not be elaborated here. Figure 8 in the embodiments shown below, which will not be elaborated here.

[0209] At this time, the video features obtained after semantic enhancement of the initial video features include the semantics corresponding to the input statement (or the semantics of the words in the input statement corresponding to the video frame), or in other words, the video features carry the semantics corresponding to the input statement (or the semantics of the words in the input statement corresponding to the video frame).

[0210] S730, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features.

[0211] Wherein, the target video segment corresponding to the input statement refers to: the content described in the input statement is included in the target video segment (or in other words, in the video frames of the target video segment).

[0212] In an embodiment of the present application, semantic enhancement is performed on a video frame according to the semantic feature to obtain the video feature of the video frame, and the semantics corresponding to the input statement can be incorporated into the video feature of the video frame. At this time, according to the semantic feature and the video feature, the target video segment corresponding to the input statement is identified, which can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0213] Optionally, the video feature of the video segment to which the video frame belongs may be determined according to the video feature, the matching degree (or similarity) between the semantic feature and the video feature of the video segment is calculated, and whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement is determined according to the matching degree.

[0214] For example, a video may be divided into multiple video clips, where a video clip may consist of one or more video frames. The video feature of the video clip is determined according to the video features of the video frames in the video clip. Subsequently, the matching degree (or similarity) between the semantic feature of the input statement and the video features of each video segment in the video is calculated, and the video segment with the largest corresponding matching degree (i.e., the value of the matching degree) is the target video segment corresponding to the input statement.

[0215] Optionally, the Euclidean distance between the semantic feature of the input statement and the video feature of the video segment may be calculated, and the matching degree between the semantic feature of the input statement and the video feature of the video segment is determined according to the calculated Euclidean distance.

[0216] Alternatively, an RNN may also be used to calculate the matching degree between the semantic feature of the input statement and the video feature of the video segment.

[0217] Optionally, before the above S730, the method 700 may further include S732.

[0218] S732, Feature fusion is performed on the video feature of the video frame using the video features of at least one other video frame to obtain the fused video feature of the video frame.

[0219] Wherein, the other video frame and the video frame belong to the same video.

[0220] In an embodiment of the present application, feature fusion is performed on the video feature of the video frame using the video features of other video frames in the video, and the context information in the video is incorporated into the video feature of the video frame, which can make the video feature of the video frame more accurate. At this time, according to the video feature, the target video segment corresponding to the input statement is identified, which can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0221] For example, the video features of the at least one other video frame can be added to the video features of the video frame, and the resulting sum is the fused video features of the video frame.

[0222] At this time, it can be considered that the video features of the at least one other video frame are fused into the fused video features of the video frame.

[0223] In the embodiments of the present application, the video features of all other video frames in the video except the video frame can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0224] Alternatively, the video features of all video frames (including the video frame) in the video can also be used to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame.

[0225] For example, the average value of the video features of all video frames (including the video frame) in the video can be calculated, and this average value is added to the video features of the video frame, and the resulting sum is the fused video features of the video frame.

[0226] For another example, the video contains a total of t video frames, and the video features of these t video frames can form the video feature sequence {f1, f2, …, f t}, where f j represents the semantic feature of the j-th word in this input statement, j is a positive integer less than t, and t is a positive integer. Multiply the video features in the video feature sequence {f1, f2, …, f t} pairwise (matrix multiplication) to obtain matrix B (matrix B can be called the correlation matrix, and the elements in matrix B can be called correlation features). Select a correlation feature for the video frame f j in the video, and add this correlation feature to the video features of the video frame f j in the video, and the resulting sum is the fused video features of the video frame f j in the video.

[0227] At this time, it can be considered that the video features of all video frames in the video are fused into the fused video features of the video frame.

[0228] Correspondingly, in S730, it is possible to determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the fused video features.

[0229] Optionally, determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic feature and the video feature may include:

[0230] Based on the video feature, determining the hierarchical structure of the video segment in the time domain; according to the semantic feature and the hierarchical structure, determining whether the video segment is the target video segment corresponding to the input statement.

[0231] It should be noted that the above video feature may be the fused video feature of the video frame obtained in S732 above, that is, based on the fused video feature of the video frame, the hierarchical structure of the video segment in the time domain can be determined.

[0232] Optionally, one-dimensional dilated convolution or one-dimensional convolution can be used to determine the hierarchical structure of the video segment in the time domain based on the video feature.

[0233] In the embodiments of the present application, using the video feature to determine the hierarchical structure of the video segment in the time domain can expand the receptive field of each video frame in the video segment while maintaining the size of the video feature of each video frame. At this time, according to the semantic feature and the hierarchical structure, identifying the target video segment corresponding to the input statement can improve the accuracy of identifying the target video segment corresponding to the input statement.

[0234] Figure 8 This is a schematic flowchart of a method for processing video in the present application. The method 800 for processing video can be executed by a device for processing video, and the device may include a feature preprocessing module 101, a feature preprocessing module 102, an integrated interaction module 103, a segment sampling module 104, and a matching degree calculation module 105. In some examples, the device for processing video may Figure 2 the execution device 110 in Figure 5 the chip shown in Figure 6 the execution device 210 in etc. devices.

[0235] Step 1:

[0236] As Figure 8 shown, the feature preprocessing module 101 preprocesses and extracts features from the input video to obtain a video feature sequence of the video; the feature preprocessing module 102 preprocesses and extracts features from the input statement to obtain a semantic feature sequence of the input statement.

[0237] Specifically, the feature preprocessing module 101 may extract features from video frames in the input video through a neural network to obtain the video feature of the video frame.

[0238] Since the video features are in vector form, this process can also be regarded as using a neural network to encode the video frames in the video into a vector.

[0239] For example, a convolutional neural network (CNN) can be used to extract features from each video frame in the input video to obtain the video features of each video frame.

[0240] After being processed by the above-mentioned feature preprocessing module 101, the video can be represented by a video feature sequence {f1, f2, …, f t}, where f j represents the video feature of the j-th video frame in the video, j is a positive integer less than t, and t is a positive integer.

[0241] Similarly, the feature preprocessing module 102 can extract features from the words in the input statement through a neural network to obtain the semantic features of the words in the input statement.

[0242] Since the semantic features are in vector form, this process can also be regarded as using a neural network to encode the words in the input statement into a vector.

[0243] For example, a bidirectional long short-term memory recurrent neural network (LSTM) can be used to extract features from each word in the input statement to obtain the semantic features of each word.

[0244] In particular, using LSTM for feature extraction can enable the semantic features of each word obtained to acquire the context information of other words (before and after each word) in the input statement.

[0245] After being processed by the above-mentioned feature preprocessing module 102, the input statement can be represented by a semantic feature sequence {w1, w2, …, w k}, where w i represents the semantic feature of the i-th word in the input statement, i is a positive integer less than k, and k is a positive integer

[0246] Step Two:

[0247] As Figure 8 shown, the integrated interaction module 103 performs integrated interaction processing on the video feature sequence of the video based on the semantic feature sequence of the input statement to obtain the candidate video feature sequence of the video.

[0248] Among them, the integrated interaction module 103 can be divided into a semantic enhancement sub-module 1031, a context interaction sub-module 1032, and a time-domain structure construction sub-module 1033. It should be noted that these sub-modules can be actual existing sub-modules or virtual modules divided according to functions, and the embodiments of the present application do not limit this.

[0249] (1) Semantic enhancement sub-module 1031

[0250] The semantic enhancement sub-module 1031 can perform semantic enhancement on the video frames in the video based on the semantic features of the input statement.

[0251] For example, the video feature sequence {f1, f2,..., f t} of the video and the semantic feature sequence {w1, w2,..., w k} of the input statement can be multiplied by a matrix to obtain a matrix A between them. The size of the matrix A is t×k, and the matrix A can be used to represent the correlation between each word in the input statement and each video frame in the video. Therefore, it can also be called a correlation matrix.

[0252] Next, the column direction of the matrix A can be normalized (for example, normalized using softmax), and then weighted processing is performed in the row direction of the matrix A. At this time, a weighted word can be selected for each video frame in the video, and a new semantic feature sequence {w`1, w`2,..., w` t} can be formed according to the semantic features corresponding to these weighted words, where w` j is the semantic feature of the weighted word corresponding to the jth video frame.

[0253] At this time, the semantic feature sequence {w`1, w`2,..., w` k} can be used as a convolution kernel to perform semantic enhancement (i.e., convolution processing) on the video frames in the video (i.e., the video feature sequence {f1, f2,..., f t}) of the video to obtain the video feature sequence of the video after semantic enhancement.

[0254] It should be noted that compared with the traditional convolution kernel, the above convolution kernel (i.e., the semantic feature sequence {w`1, w`2,..., w` k}) has two main differences:

[0255] 1. The weights of the convolution kernel are not included in the model but are dynamically determined by the input semantics.

[0256] Dynamically determining the weights of the convolutional kernels by the input statement can make the model very flexible. The video features of the extracted video frames can be determined by the semantics of the input statement, and it is very convenient to detect other interesting video segments in the same video by replacing the input statement.

[0257] 2. Convolution processing usually uses the same convolutional kernel for transformation at each position, while in the semantic enhancement sub-module 1031, each video frame is convolved (i.e., semantically enhanced) by its corresponding weighted word.

[0258] Convolving each video frame with its corresponding weighted word can explore the relationship between the video frame and the word at a finer granularity. Using the details corresponding to the word (the semantic features corresponding to the word) to semantically enhance the video frame can make the video features of the video frame more accurate.

[0259] (2) Context interaction sub-module 1032

[0260] The context interaction sub-module 1032 can incorporate the content of other video frames in the video (i.e., context information) into the video frames in the video.

[0261] Optionally, in the context interaction sub-module 1032, context interaction can be performed in the following two ways, specifically as follows:

[0262] Method 1:

[0263] Perform context interaction in the way of average pooling.

[0264] For example, for the video feature sequence {f1, f2,..., f t} of the video, calculate the average value of all video features in this video feature sequence to obtain the average value f`. Subsequently, add the average value f` to the video feature f j of the j-th video frame in the video.

[0265] Method 2:

[0266] Another way is similar to the way of semantic enhancement in the semantic enhancement sub-module 1031. This way is similar to the perception method of two different modalities in the 1031 module.

[0267] For example, for the video feature sequence {f1, f2,..., f t} of the video, for this video feature sequence {f1, f2,..., f tMultiply the video features in {} pairwise (matrix multiplication) to obtain matrix B. The size of matrix B is t×t. Next, the columns of matrix B can be normalized (for example, normalized using softmax), and then weighted processing is performed in the row direction of matrix B. At this time, a weighted video frame can be selected for each video frame in the video, and the video features of the weighted video frame are added to each video frame, that is, the context interaction of each video frame is completed.

[0268] (3) Temporal Structure Construction Sub-module 1033

[0269] The temporal structure construction sub-module 1033 can construct the hierarchical structure of the time-frequency in the time domain.

[0270] Optionally, the temporal structure construction sub-module 1033 can receive the video feature sequence processed by the context interaction sub-module 1032, perform one-dimensional dilated convolution on the video feature sequence to obtain the hierarchical structure of the video in the time domain (i.e., the candidate video feature sequence of the video).

[0271] Step Three:

[0272] As Figure 8 shown, the segment sampling module 104 samples the candidate video feature sequence of the video to obtain the video feature sequences of multiple video segments.

[0273] Optionally, the segment sampling module 104 receives the candidate video feature sequence of the video and generates multiple video segments of the video based on a preset rule.

[0274] For example, in the candidate video feature sequence of the video, the video features of 7 video frames can be sampled at equal intervals in chronological order. Then, the video features of these 7 video frames can form a video segment in the chronological order between them, that is, a video segment is generated.

[0275] Among them, in the case where the time coordinates of the sampled video frames are not aligned with the time coordinates of the video frames in the video, the method of linear interpolation can be used to align the time coordinates of the sampled video frames.

[0276] Optionally, the video features of the sampled video segments can be input into the integrated interaction module 103 for integrated interaction processing.

[0277] Step Four:

[0278] As Figure 8As shown, the integrated interaction module 103 performs integrated interaction processing on the video feature sequences of the multiple video segments based on the semantic feature sequence of the input statement, and obtains the candidate video feature sequences of the multiple video segments.

[0279] Optionally, in step four, the method for the integrated interaction module 103 to perform integrated interaction processing on the video feature sequences of the multiple video segments is similar to that in step two above, and will not be elaborated here.

[0280] Among them, when the time-domain structure construction sub-module 1033 determines the hierarchical structure of the video segment in the time domain (that is, the candidate video feature sequence of the video segment), one-dimensional convolution can be performed on the video feature sequence of the video segment.

[0281] Step five:

[0282] As Figure 8 shown, the matching degree calculation module 105 calculates the matching degree (or similarity) between the semantic feature sequence of the input statement and the video feature sequences of each of the multiple video segments, and determines whether the video segment (corresponding to the matching degree) is the target video segment corresponding to the input statement according to the calculated matching degree.

[0283] For example, the video segment with the largest matching degree (the value of the matching degree) among the multiple video segments can be determined as the target video segment.

[0284] It should be noted that the target video segment here can refer to the video segment among the multiple video segments generated by the video that best matches the semantic features of the input statement; or it can also refer to the content in the video segment that is most similar (or closest) to the semantics expressed by the input statement.

[0285] To illustrate the effect of the method for processing videos in this application embodiment, the accuracy rate of identifying the target video segment corresponding to the input statement by the method for processing videos based on this application embodiment will be analyzed below in combination with specific test results.

[0286] Table 1

[0287] Method Rank@1 Accuracy TMN 22.92% MCN 28.10% TGN 28.23% This application 32.45%

[0288] Table 1 shows the accuracy rates of identifying the target video segment corresponding to the input statement using different schemes on the DiDeMo dataset.

[0289] As can be seen from Table 1, the accuracy rate of recognition using the method in the temporal modular network (TMN) is 22.92%, the accuracy rate of recognition using the method in the moment context network (MCN) is 28.10%, the accuracy rate of recognition using the method in the temporal ground net (TGN) is 28.23%, and the accuracy rate of recognition using the method for processing videos in this application is 32.45%. It can be seen that compared with several methods in Table 1, the method for processing videos in this application can have a significant improvement in the accuracy of Rank@1.

[0290] Table 2

[0291] Method IoU = 0.5, Rank@1 Accuracy IoU = 0.7, Rank@1 Accuracy CTRL 23.63% 8.89% ACL 30.48% 12.20% SAP 27.42% 13.36% LSTM 35.6% 15.8% This application 41.69% 22.88%

[0292] Table 2 shows the accuracy rates of recognizing the target video segment corresponding to the input statement using different schemes on the Charades-STA dataset, where IoU is the intersection over union (IoU).

[0293] As can be seen from Table 2, when IoU = 0.5, the accuracy rate of recognition using the method in the cross-modal temporal regression localizer (CTRL) is 23.63%, the accuracy rate of recognition using the method in the activity concept based localizer (ACL) is 30.48%, the accuracy rate of recognition using the method in the semantic activity proposal (SAP) is 27.42%, the accuracy rate of recognition using the method in the long-short term memory (LSTM) is 35.6%, and the accuracy rate of recognition using the method for processing videos in this application is 41.69%.

[0294] When IoU = 0.7, the accuracy rate of recognition using the method in CTRL is 8.89%, the accuracy rate of recognition using the method in ACL is 12.20%, the accuracy rate of recognition using the method in SAP is 13.36%, the accuracy rate of recognition using the method in LSTM is 15.8%, and the accuracy rate of recognition using the method for processing videos in this application is 22.88%.

[0295] It can be seen that compared with several methods in Table 2, the method for processing videos in this application can achieve a significant improvement in the accuracy of Rank@1.

[0296] In summary, in the embodiment of this application, semantic enhancement is performed on the video frame according to the semantic feature to obtain the video feature of the video frame, and the semantics corresponding to the input statement can be incorporated into the video feature of the video frame. At this time, according to the semantic feature and the video feature, the target video segment corresponding to the input statement is recognized, which can effectively improve the accuracy of recognizing the target video segment corresponding to the input statement.

[0297] Figure 9 It is a schematic diagram of the hardware structure of the device for processing videos in the embodiment of this application. Figure 9 The shown device 4000 for processing videos includes a memory 4001, a processor 4002, a communication interface 4003, and a bus 4004. Among them, the memory 4001, the processor 4002, and the communication interface 4003 are communicatively connected to each other through the bus 4004.

[0298] The memory 4001 can be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 4001 can store a program. When the program stored in the memory 4001 is executed by the processor 4002, the processor 4002 and the communication interface 4003 are used to execute each step of the device for processing videos in the embodiment of this application.

[0299] The processor 4002 can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), or one or more integrated circuits, and is used to execute relevant programs to implement the functions required by the units in the device for processing videos in the embodiment of this application, or to execute the method for processing videos in the method embodiment of this application.

[0300] The processor 4002 can also be an integrated circuit chip with signal processing capabilities. During the implementation process, each step of the method for processing videos in the embodiment of this application can be completed by the integrated logic circuit in the hardware of the processor 4002 or by instructions in software form.

[0301] The above-mentioned processor 4002 may also be a general-purpose processor, a digital signal processor (DSP), an ASIC, a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The above-mentioned general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 4001, and the processor 4002 reads the information in the memory 4001 and combines its hardware to complete the functions required to be executed by the units included in the video processing device in the embodiments of the present application, or executes the video processing method in the method embodiments of the present application.

[0302] The communication interface 4003 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the device 4000 and other devices or communication networks. For example, input statements and video frames (or videos) to be processed can be obtained through the communication interface 4003.

[0303] The bus 4004 may include a path for transmitting information between various components of the device 4000 (for example, the memory 4001, the processor 4002, and the communication interface 4003).

[0304] Figure 10 It is a schematic diagram of the hardware structure of the model training device 5000 in the embodiments of the present application. Similar to the above-mentioned device 4000, Figure 10 The shown model training device 5000 includes a memory 5001, a processor 5002, a communication interface 5003, and a bus 5004. Among them, the memory 5001, the processor 5002, and the communication interface 5003 are communicatively connected to each other through the bus 5004.

[0305] The memory 5001 may store a program. When the program stored in the memory 5001 is executed by the processor 5002, the processor 5002 is used to execute each step of the training method for training the video processing device in the embodiments of the present application.

[0306] The processor 5002 may adopt a general-purpose CPU, a microprocessor, an ASIC, a GPU, or one or more integrated circuits, and is used to execute relevant programs to implement the training method for training the video processing device in the embodiments of the present application.

[0307] The processor 5002 can also be an integrated circuit chip with the ability to process signals. During the implementation of the training process, each step of the training method of the device for processing videos according to the embodiments of the present application can be completed by the integrated logic circuit of the hardware in the processor 5002 or instructions in the form of software.

[0308] It should be understood that by Figure 10 training the device for processing videos through the model training device 5000 shown, the trained device for processing videos can be used to execute the method for processing videos according to the embodiments of the present application. Specifically, by training the neural network through the device 5000, the device for processing videos in the Figure 5 method shown can be obtained, or Figure 6 the device for processing videos shown.

[0309] Specifically, Figure 10 the device shown can obtain training data and the device for processing videos to be trained from the outside through the communication interface 5003, and then the processor trains the device for processing videos to be trained according to the training data.

[0310] Optionally, the above training data may include input statements, training videos, and video segments in the training videos that have the highest matching degree with the input statements. Among them, the video segments in the training videos that have the highest matching degree with the input statements can be video segments pre-annotated manually.

[0311] It should be noted that although the above devices 4000 and 5000 only show a memory, a processor, and a communication interface, in the specific implementation process, those skilled in the art should understand that the devices 4000 and 5000 may also include other devices necessary for normal operation. At the same time, according to specific needs, those skilled in the art should understand that the devices 4000 and 5000 may also include hardware devices for implementing other additional functions. In addition, those skilled in the art should understand that the devices 4000 and 5000 may also only include the devices necessary for implementing the embodiments of the present application, and do not necessarily include Figure 9 and Figure 10 all the devices shown in.

[0312] It should be understood that the processor in the embodiments of the present application may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0313] It should also be understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of random access memory (RAM) are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced SDRAM (ESDRAM), synchlink DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0314] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0315] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this document generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.

[0316] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0317] It should be understood that in various embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0318] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0319] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0320] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.

[0321] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0322] In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0323] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0324] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

Claims

1. A method for processing video, characterized in that, Including: Obtain the semantic features of the input statement; Semantically enhance the video frame according to the semantic features to obtain the video features of the video frame, where the semantic features are included in the video features. Among them, the video features are obtained through convolution processing, and the convolution kernel of the convolution processing is determined according to the semantic features; Determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features.

2. The method according to claim 1, characterized in that, The semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame includes: Determine the word in the input statement corresponding to the video frame; Semantically enhance the video frame according to the semantic features of the word corresponding to the video frame to obtain the video features of the video frame.

3. The method according to claim 1 or 2, characterized in that, The semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame in the video includes: Extract features from the video frame according to the semantic features to obtain the video features of the video frame.

4. The method according to claim 1 or 2, characterized in that, The method further includes: Obtain the initial video features of the video frame; Among them, the semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame includes: Semantically enhance the initial video features according to the semantic features to obtain the video features of the video frame.

5. The method according to any one of claims 1 to 4, characterized in that, The method further includes: Use the video features of at least one other video frame to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame. The other video frame and the video frame belong to the same video; Among them, the determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features includes: Determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the fused video features.

6. The method according to any one of claims 1 to 5, characterized in that, The determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features includes: Based on the video features, determine the hierarchical structure of the video segment in the time domain; Determine whether the video segment is the target video segment corresponding to the input statement according to the semantic features and the hierarchical structure.

7. An apparatus for processing video, characterized in that, Including: Obtain the semantic features of the input statement; Semantically enhance the video frame according to the semantic features to obtain the video features of the video frame, where the semantic features are included in the video features. Among them, the video features are obtained through convolution processing, and the convolution kernel of the convolution processing is determined according to the semantic features; Determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features.

8. The apparatus according to claim 7, characterized in that, The semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame includes: Determine the word in the input statement corresponding to the video frame; Semantically enhance the video frame according to the semantic features of the word corresponding to the video frame to obtain the video features of the video frame.

9. The apparatus according to claim 7 or 8, characterized in that, The semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame in the video includes: Extract features from the video frame according to the semantic features to obtain the video features of the video frame.

10. The apparatus according to claim 7 or 8, characterized in that, Obtain the initial video features of the video frame; wherein, the semantically enhancing the video frame according to the semantic features to obtain the video features of the video frame includes: Semantically enhance the initial video features according to the semantic features to obtain the video features of the video frame.

11. The device according to any one of claims 7 to 10, characterized in that, Use the video features of at least one other video frame to perform feature fusion on the video features of the video frame to obtain the fused video features of the video frame, where the other video frame and the video frame belong to the same video; wherein, the determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features includes: Determine whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the fused video features.

12. The device according to any one of claims 7 to 11, characterized in that, The determining whether the video segment to which the video frame belongs is the target video segment corresponding to the input statement according to the semantic features and the video features includes: Based on the video features, determine the hierarchical structure of the video segment in the time domain; According to the semantic features and the hierarchical structure, determine whether the video segment is the target video segment corresponding to the input statement.

13. A device for processing video, characterized in that, Comprising a processor and a memory, the memory is used for storing program instructions, and the processor is used for calling the program instructions to execute the method according to any one of claims 1 to 6.

14. A computer-readable storage medium, characterized in that, The computer-readable medium stores program code for a device to execute, and the program code includes for executing the method according to any one of claims 1 to 6.

15. A chip, characterized in that, The chip includes a processor and a data interface, and the processor reads the instructions stored on the memory through the data interface to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video positioning method and device and electronic equipment

    CN110225368A

  • Cross-modal video moment retrieval method based on cross-modal dynamic convolutional network

    CN112650886A

  • System and method for localization of activities in videos

    US10839223B1