Large-scale video behavior recognition method based on Fourier transform saliency

The frequency domain transformation matrix and mask processing are generated through the Fourier transform significance method. Combined with the behavior recognition model, the low accuracy and resource waste caused by background interference in video behavior recognition are solved, and more efficient video behavior recognition is achieved.

CN120298957AActive Publication Date: 2025-07-11HANGZHOU INNOVATION RES INST OF BEIJING UNIV OF AERONAUTICS & ASTRONAUTICS +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510774034.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-07-11
Estimated Expiration
2045-06-11

AI Technical Summary

Technical Problem

Existing video behavior recognition methods are susceptible to background information interference, resulting in low recognition accuracy and waste of computing resources.

Method used

Using a method based on Fourier transform significance, a frequency domain transformation matrix is generated, the target video frame is determined and masked, and feature extraction and recognition are performed in combination with a pre-trained behavior recognition model.

Benefits of technology

The accuracy of behavior recognition is improved, the waste of computing resources is reduced, and the need for processing per frame is reduced through the processing of significant video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298957A_ABST
    Figure CN120298957A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a large-scale video behavior recognition method based on Fourier transform saliency. A specific embodiment of the method comprises the following steps: generating a to-be-processed video frame group corresponding to to-be-processed video data; generating each frequency domain transformation matrix corresponding to the to-be-processed video frame group; determining a target video frame corresponding to the to-be-processed video data; generating target problem text data; performing mask processing on the target problem text data to obtain mask problem text data; inputting the target video frame into an image processing layer in a pre-trained behavior recognition model; inputting the mask problem text data into a text processing layer in the behavior recognition model; inputting the video frame feature information and the video text feature information into an output layer in a behavior recognition model to obtain behavior description information; and sending the behavior description information to the user terminal. According to the embodiment, computing resources consumed during identification are reduced, and waste of the computing resources can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of computer technologies, and particularly to a method for large-scale video behavior recognition based on Fourier transform saliency. Background Art

[0002] Video behavior recognition is a technology for recognizing behaviors in videos. Currently, when recognizing behaviors in videos, the commonly adopted method is as follows: First, feature extraction is performed on each video frame in the video through a feature extraction algorithm to obtain a feature vector corresponding to the video. Then, the feature vector is classified through a pre-trained classifier to obtain behavior description information corresponding to the video.

[0003] However, when recognizing behaviors in videos in the above manner, the following technical problems often exist: When recognizing behaviors in videos, the recognition process is easily interfered by the background information in the videos, resulting in a low recognition accuracy. As a result, the probability of wasting computing resources due to the need to re-recognize the video is relatively high, leading to waste of computing resources. Moreover, when using a feature extraction algorithm to extract features from a video, each frame in the video needs to be processed, resulting in a large amount of computing resources consumed during processing. Summary of the Invention

[0004] The content part of the present disclosure is used to introduce concepts in a brief form, and these concepts will be described in detail in the following detailed implementation part. The content part of the present disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0005] Some embodiments of the present disclosure propose a method for large-scale video behavior recognition based on Fourier transform saliency to solve one or more of the technical problems mentioned in the above background art part.

[0006] In a first aspect, some embodiments of the present disclosure provide a method for large-scale video behavior recognition based on Fourier transform saliency. The method includes: in response to receiving the video data to be processed and the question text data sent by the user terminal, generating a group of video frames to be processed corresponding to the video data to be processed; based on the group of video frames to be processed, generating respective frequency domain transformation matrices corresponding to the group of video frames to be processed; based on the respective frequency domain transformation matrices, determining a target video frame corresponding to the video data to be processed; based on the question text data, generating target question text data; performing a masking process on the target question text data to obtain masked question text data; inputting the target video frame into an image processing layer in a pre-trained behavior recognition model to obtain video frame feature information, wherein the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; inputting the masked question text data into a text processing layer in the behavior recognition model to obtain video text feature information, wherein the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; inputting the video frame feature information and the video text feature information into an output layer in the behavior recognition model to obtain behavior description information; and sending the behavior description information to the user terminal.

[0007] In a second aspect, some embodiments of the present disclosure provide a device for large-scale video behavior recognition based on Fourier transform saliency. The device includes: a first generation unit configured to generate a group of video frames to be processed corresponding to the video data to be processed in response to receiving the video data to be processed and the question text data sent by the user terminal; a second generation unit configured to generate respective frequency domain transformation matrices corresponding to the group of video frames to be processed based on the group of video frames to be processed; a determination unit configured to determine a target video frame corresponding to the video data to be processed based on the respective frequency domain transformation matrices; a third generation unit configured to generate target question text data based on the question text data; a masking unit configured to perform a masking process on the target question text data to obtain masked question text data; a first input unit configured to input the target video frame into an image processing layer in a pre-trained behavior recognition model to obtain video frame feature information, wherein the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; a second input unit configured to input the masked question text data into a text processing layer in the behavior recognition model to obtain video text feature information, wherein the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; a third input unit configured to input the video frame feature information and the video text feature information into an output layer in the behavior recognition model to obtain behavior description information; and a sending unit configured to send the behavior description information to the user terminal.

[0008] In a third aspect, some embodiments of the present disclosure provide an electronic device, including: one or more processors; a storage device storing thereon one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the method described in any implementation manner of the first aspect above.

[0009] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium storing a computer program thereon, wherein the program, when executed by a processor, implements the method described in any implementation manner of the first aspect above.

[0010] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the large-scale video behavior recognition method based on Fourier transform significance in some embodiments of the present disclosure, waste of computing resources can be reduced. Specifically, the reason for the waste of computing resources is that when recognizing the behavior in a video, the recognition process is easily interfered by the background information in the video, resulting in a low recognition accuracy. As a result, the probability of consuming computing resources to recognize the video again is relatively high, which in turn leads to waste of computing resources. Moreover, when using a feature extraction algorithm to extract features from a video, each frame in the video needs to be processed, resulting in a large amount of computing resources consumed during processing. Based on this, in some embodiments of the large-scale video behavior recognition method based on Fourier transform significance of the present disclosure, first, in response to receiving the video data to be processed and the problem text data sent by the user terminal, a group of video frames to be processed corresponding to the video data to be processed is generated. Thus, the video data to be recognized can be obtained. Secondly, based on the group of video frames to be processed, various frequency domain transformation matrices corresponding to the group of video frames to be processed are generated. Thus, by performing Fourier transform processing on the video, the model can better recognize foreground information during recognition and reduce the interference of background information. Then, based on the various frequency domain transformation matrices, the target video frame corresponding to the video data to be processed is determined. Thus, the video frame to be processed with the highest significance can be determined from the group of video frames to be processed. Then, based on the problem text data, the target problem text data is generated. Then, the target problem text data is masked to obtain the masked problem text data. Thus, the redundant information in the text data can be masked. Then, the target video frame is input into the image processing layer of the pre-trained behavior recognition model to obtain video frame feature information, where the image processing layer includes an image mapping network, an image embedding network, and an image linear layer. Thus, the target video frame can be processed by the behavior recognition model. Then, the masked problem text data is input into the text processing layer of the behavior recognition model to obtain video text feature information, where the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer. Thus, the masked problem text data can be processed by the behavior recognition model. Then, the video frame feature information and the video text feature information are input into the output layer of the behavior recognition model to obtain behavior description information. Thus, the behavior description information corresponding to the video can be obtained. Finally, the behavior description information is sent to the user terminal. Also, because a most significant video frame to be processed can be determined from the group of video frames to be processed of the video data to be processed first, and then the most significant video frame is processed to obtain the behavior description information, instead of processing each video frame to be processed in the video data to be processed, the computing resources consumed during processing can be reduced.It is also because the video can be processed by Fourier transform first, enabling the model to better process foreground information and reducing the interference of background information on recognition. Therefore, the recognition accuracy can be improved, and further, the probability of the situation where computational resources need to be consumed to re-recognize the video can be reduced, and further, the waste of computational resources can be reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.

[0012] Figure 1 is a flowchart of some embodiments of a large-scale video behavior recognition method based on Fourier transform saliency according to the present disclosure; Figure 2 is a schematic structural diagram of some embodiments of a large-scale video behavior recognition device based on Fourier transform saliency according to the present disclosure; Figure 3 is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0014] In addition, it should be noted that only parts related to the relevant invention are shown in the drawings for the sake of convenience of description. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0015] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.

[0016] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".

[0017] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are for illustrative purposes only and are not used to limit the scope of these messages or information.

[0018] The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments.

[0019] Figure 1 Flow 100 of some embodiments of a large-scale video behavior recognition method based on Fourier transform saliency according to the present disclosure is shown. The large-scale video behavior recognition method based on Fourier transform saliency includes the following steps: Step 101, in response to receiving the video data to be processed and the question text data sent by the user terminal, generate a group of video frames to be processed corresponding to the video data to be processed.

[0020] In some embodiments, the execution subject (such as a computing device) of the large-scale video behavior recognition method based on Fourier transform saliency may, in response to receiving the video data to be processed and the question text data sent by the user terminal, generate a group of video frames to be processed corresponding to the above-mentioned video data to be processed. Among them, the above-mentioned video data to be processed may be the video data to be recognized. The above-mentioned question text data may be a question raised by the user according to the above-mentioned video data to be processed. For example, the above-mentioned question text data may be "What is the person in the video doing". Each video frame to be processed in the above-mentioned group of video frames to be processed may be a video frame in the above-mentioned video data to be processed. Each video frame to be processed in the above-mentioned group of video frames to be processed may be an RGB image. Each pixel point in each pixel point included in the video frame to be processed corresponds to a red color channel value, a green color channel value, and a blue color channel value. The above-mentioned red color channel value may be the color channel value corresponding to the R channel in the RGB image. The above-mentioned green color channel value may be the color channel value corresponding to the G channel in the RGB image. The above-mentioned blue color channel value may be the color channel value corresponding to the B channel in the RGB image. Each video frame to be processed in the above-mentioned group of video frames to be processed corresponds to image height data and image width data. The above-mentioned image height data may be the image height corresponding to the video frame to be processed. The above-mentioned image width data may be the image width corresponding to the video frame to be processed. The above-mentioned execution subject may be a server. In practice, the above-mentioned execution subject may determine each video frame in the above-mentioned video data to be processed as a group of video frames to be processed.

[0021] Step 102, based on the group of video frames to be processed, generate respective frequency domain transformation matrices corresponding to the group of video frames to be processed.

[0022] In some embodiments, the above-mentioned execution entity may generate respective frequency-domain transformation matrices corresponding to the to-be-processed video frame group based on the to-be-processed video frame group. Each of the frequency-domain transformation matrices among the respective frequency-domain transformation matrices may be a complex matrix used to represent the frequency domain of the corresponding to-be-processed video frame.

[0023] In some alternative implementation manners of some embodiments, the above-mentioned execution entity may generate respective frequency-domain transformation matrices corresponding to the to-be-processed video frame group based on the to-be-processed video frame group through the following steps: For each to-be-processed video frame in the to-be-processed video frame group, perform the following steps: First step, perform image conversion processing on the to-be-processed video frame based on a preset first gray coefficient value, a preset second gray coefficient value, and a preset third gray coefficient value to obtain gray video frame data. The first gray coefficient value may be a value used to convert the red color channel value of each element in the to-be-processed video frame. For example, the first gray coefficient value may be 0.299. The second gray coefficient value may be a value used to convert the green color channel value of each element in the to-be-processed video frame. For example, the second gray coefficient value may be 0.587. The third gray coefficient value may be a value used to convert the blue color channel value of each element in the to-be-processed video frame. For example, the blue color channel value may be 0.114. The gray video frame data may be a grayscale image corresponding to the to-be-processed video frame. In practice, for each pixel point among the respective pixel points included in the to-be-processed video frame, first, the execution entity may determine the product of the red color channel value corresponding to the pixel point and the first gray coefficient value as the first channel value. Then, the product of the green color channel value corresponding to the pixel point and the second gray coefficient value may be determined as the second channel value. Then, the product of the blue color channel value corresponding to the pixel point and the third gray coefficient value may be determined as the third channel value. Then, the sum of the first channel value, the second channel value, and the third channel value may be determined as the gray value corresponding to the pixel point. Then, according to the positions corresponding to the respective pixel points in the to-be-processed video frame, the determined gray values may be combined into a matrix as a gray matrix. Finally, the gray matrix may be rendered through an image rendering function to obtain the gray video frame data. The image rendering function may be a function capable of rendering pixel values into an image. For example, the image rendering function may be an imshow function.

[0024] Second step, perform frequency-domain transformation processing on the gray video frame data to obtain a frequency-domain transformation matrix.

[0025] In some alternative implementations of some embodiments, the above-mentioned execution entity may perform frequency-domain transformation processing on the above-mentioned grayscale video frame data through the following steps to obtain a frequency-domain transformation matrix: First, based on the image height data and image width data corresponding to the video frame to be processed, determine each matrix height data and each matrix width data. Each matrix height data among the above-mentioned matrix height data may be a non-negative integer less than the above-mentioned image height data. Each matrix width data among the above-mentioned matrix width data may be a non-negative integer less than the above-mentioned image width data. In practice, the above-mentioned execution entity may determine each non-negative integer less than the above-mentioned image height data as each matrix height data. Each non-negative integer less than the above-mentioned image width data may be determined as each matrix width data.

[0026] Second, based on the above-mentioned grayscale video frame data, determine a grayscale video matrix. The above-mentioned grayscale video matrix may be an image matrix corresponding to the above-mentioned grayscale video frame data. In practice, the above-mentioned execution entity may determine the image matrix corresponding to the above-mentioned grayscale video frame data as the grayscale video matrix.

[0027] Third, for each matrix height data among the above-mentioned matrix height data and each matrix width data among the above-mentioned matrix width data, perform the following steps: The first sub-step: For each element among the elements included in the above-mentioned grayscale video matrix, perform the following steps: Sub-step one: Based on the above-mentioned grayscale video matrix and the above-mentioned element, determine the element height data and element width data corresponding to the above-mentioned element. The above-mentioned element height data may be the height corresponding to the above-mentioned element in the above-mentioned grayscale video matrix. The above-mentioned element width data may be the width corresponding to the above-mentioned element in the above-mentioned grayscale video matrix. In practice, the above-mentioned execution entity may determine the image height of the above-mentioned element in the above-mentioned grayscale video matrix as the element height. Then, the difference between the above-mentioned element height and a preset value may be determined as the element height data. Then, the image width corresponding to the above-mentioned element in the above-mentioned grayscale video matrix may be determined as the element width. Then, the difference between the above-mentioned element width and the above-mentioned preset value may be determined as the element width data. The above-mentioned preset value may be 1.

[0028] Sub-step two: Determine the product of the above-mentioned matrix height data and the above-mentioned element height data as the matrix element height data.

[0029] Sub-step three: Determine the product of the above-mentioned matrix width data and the above-mentioned element width data as the matrix element width data.

[0030] Sub-step four, determine the height ratio data by taking the ratio of the above matrix element height data to the image height data corresponding to the video frame to be processed.

[0031] Sub-step five, determine the width ratio data by taking the ratio of the above matrix element width data to the image width data corresponding to the video frame to be processed.

[0032] Sub-step six, determine the sum of the above height ratio data and the above width ratio data as the sum ratio data.

[0033] Sub-step seven, determine the exponential data by taking the product of the preset imaginary data, the preset coefficient data, and the above sum ratio data. Among them, the above imaginary data can be the imaginary unit. The above coefficient data can be the product of pi and a preset multiple. The above preset multiple can be a preset value. For example, the above preset multiple can be 2.

[0034] Sub-step eight, generate the exponential data to be processed based on the preset exponential base data and the above exponential data. Among them, the above exponential base data can be the natural constant e. In practice, the above execution entity can determine the exponential data to be processed as the power of the above exponential base data to the above exponential data. For example, when the above exponential base data is the natural constant e and the above exponential data is 2πi, the above exponential data to be processed can be e to the power of 2πi. Among them, i can represent the above imaginary data.

[0035] Sub-step nine, determine the matrix sub-element data by taking the product of the element value corresponding to the above element and the above exponential data to be processed.

[0036] The second sub-step, determine the sum of the determined matrix sub-element data as the matrix element data corresponding to the above matrix height data and the above matrix width data.

[0037] Fourth step, combine the obtained matrix element data into a frequency domain transformation matrix. In practice, the above execution entity can combine the above matrix element data into a frequency domain transformation matrix according to the matrix height data and the matrix width data corresponding to each matrix element data among the above matrix element data. For example, when the matrix height data corresponding to the matrix element data is 3 and the corresponding matrix width data is 2, the matrix element data should be arranged at the position of the 3rd row and the 2nd column in the frequency domain transformation matrix.

[0038] Step 103, determine the target video frame corresponding to the video data to be processed based on each frequency domain transformation matrix.

[0039] In some embodiments, the above-mentioned execution entity may determine a target video frame corresponding to the to-be-processed video data based on the above-mentioned various frequency-domain transformation matrices. Wherein, the above-mentioned target video frame may be a to-be-processed video frame selected from the above-mentioned to-be-processed video frame group for action recognition.

[0040] In some alternative implementation manners of some embodiments, the above-mentioned execution entity may determine a target video frame corresponding to the to-be-processed video data based on the above-mentioned various frequency-domain transformation matrices through the following steps: First step, for each frequency-domain transformation matrix in the above-mentioned various frequency-domain transformation matrices, perform the following steps: First sub-step, generate a spectral energy matrix based on the above-mentioned frequency-domain transformation matrix. Wherein, the above-mentioned spectral energy matrix may be a real number matrix used to characterize the energy distribution of the corresponding to-be-processed video frame in the frequency domain. In practice, first, the above-mentioned execution entity may input the above-mentioned frequency-domain transformation matrix into an absolute value function, and obtain the matrix output by the above-mentioned absolute value function as the amplitude matrix. Wherein, the above-mentioned absolute value function may be a function capable of generating the amplitude spectrum of a complex matrix. For example, the above-mentioned absolute value function may be the abs function. Then, the logarithm of the above-mentioned amplitude matrix may be determined as the amplitude logarithm matrix. Finally, the product of a preset energy coefficient and the above-mentioned amplitude logarithm matrix may be determined as the spectral energy matrix. Wherein, the above-mentioned energy coefficient may be a preset value. Here, the specific setting of the above-mentioned energy coefficient is not limited. For example, the above-mentioned energy coefficient may be 0.9.

[0041] Second sub-step, generate spectral element accumulation data based on the above-mentioned spectral energy matrix. Wherein, the above-mentioned spectral element accumulation data may be the average value corresponding to each element value in the above-mentioned spectral energy matrix. In practice, for each element value included in the above-mentioned spectral energy matrix, the above-mentioned execution entity may determine the average value corresponding to each element value as the spectral element accumulation data.

[0042] Third sub-step, determine the product of the image height data and the image width data corresponding to the above-mentioned to-be-processed video frame as the image processing data.

[0043] Fourth sub-step, determine the ratio of the above-mentioned spectral element accumulation data to the above-mentioned image processing data as the average energy data.

[0044] Second step, determine the average energy data that satisfies the preset energy data condition among the determined average energy data as the target average energy data. Wherein, the above-mentioned preset energy data condition may be that the value of the average energy data is the largest among the above-mentioned average energy data.

[0045] Third step, determine the to-be-processed video frame corresponding to the above-mentioned target average energy data in the above-mentioned to-be-processed video frame group as the target video frame.

[0046] Step 104: Generate target question text data based on the question text data.

[0047] In some embodiments, the above-mentioned execution subject may generate target question text data based on the above-mentioned question text data. The above-mentioned target question text data may be text data after the above-mentioned question text data is sorted. In practice, the above-mentioned execution subject may supplement the above-mentioned question text data with a preset user text to obtain the target question text data. The above-mentioned user text may be "Questioner: User". As an example, when the question text data is "What are the people in the video doing" and the user text is "Questioner: User", the generated target question text data may be "What are the people in the video doing; Questioner: User".

[0048] Step 105, masking the target question text data to obtain masked question text data.

[0049] In some embodiments, the execution subject may perform masking on the target question text data to obtain masked question text data. The masked question text data may be text data obtained by masking some contents in the target question text data. In practice, first, the execution subject may use a search instruction to search for each character in the target question text data that satisfies a preset replacement condition as each character to be replaced. Then, a replacement instruction may be used to replace each character to be replaced in the target question text data with a preset mask character to obtain masked question text data. The mask character may be [MASK]. The preset replacement condition may be that the character exists in a preset replacement character group. Each replacement character in the replacement character group may be a character pre-set by a technician and needs to be replaced with the mask character. For example, the replacement character group may include but is not limited to: "this", "that", "what".

[0050] Step 106: input the target video frame into the image processing layer of the pre-trained behavior recognition model to obtain video frame feature information.

[0051] In some embodiments, the execution subject may input the target video frame into the image processing layer of a pre-trained behavior recognition model to obtain video frame feature information. The video frame feature information may be a linearly transformed feature vector corresponding to the target video frame. The behavior recognition model may be a neural network model that takes the target video frame and masked question text data as input and takes behavior description information as output. The behavior recognition model may include three layers.

[0052] The first layer can be an image processing layer. Among them, the above-mentioned image processing layer can include an image mapping network, an image embedding network, and an image linear layer.

[0053] The above-mentioned image mapping network can be a neural network that takes a target video frame as input and outputs video frame information corresponding to the target video frame. For example, the above-mentioned image mapping network can be a convolutional neural network. Among them, the above-mentioned video frame information can be a feature vector corresponding to the target video frame.

[0054] The above-mentioned image embedding network can be a neural network that takes the video frame information corresponding to the target video frame as input and outputs a weighted matrix corresponding to the target video frame. For example, the above-mentioned neural network can be a Transformer. The above-mentioned image embedding network corresponds to a query weight matrix, a key weight matrix, and a value weight matrix.

[0055] In practice, first, the above-mentioned image embedding network can determine the product of the query weight matrix and the above-mentioned video frame information as the query matrix. Second, the above-mentioned image embedding network can determine the product of the key weight matrix and the above-mentioned video frame information as the key matrix. Then, the above-mentioned image embedding network can determine the product of the value weight matrix and the above-mentioned video frame information as the value matrix. Then, the above-mentioned image embedding network can determine the transpose of the above-mentioned key matrix as the key transpose matrix. Then, the product of the above-mentioned query matrix and the above-mentioned key transpose matrix can be determined as the product matrix. Then, the dimension of the above-mentioned video frame information can be determined as the feature dimension data. Then, the square root of the above-mentioned feature dimension data can be determined as the dimension square root data. Then, the ratio of the above-mentioned product matrix to the above-mentioned dimension square root data can be determined as the dimension ratio data. Then, the above-mentioned dimension ratio data can be input into a normalization function, and the data output by the above-mentioned normalization function is used as the attention score data. Finally, the product of the above-mentioned attention score data and the above-mentioned value matrix can be determined as the weighted matrix. Among them, the above-mentioned query weight matrix can be a weight matrix for generating the query matrix. The above-mentioned key weight matrix can be a weight matrix for generating the key matrix. The above-mentioned value weight matrix can be a weight matrix for generating the value matrix. The above-mentioned normalization function can be a Softmax function.

[0056] The above-mentioned image linear layer can be a linear layer that takes the weighted matrix corresponding to the target video frame as input and outputs video frame feature information corresponding to the target video frame. The above-mentioned image linear layer corresponds to a linear matrix and an image bias vector.

[0057] In practice, first, the above-mentioned image linear layer can determine the product of the linear matrix and the above-mentioned weighted matrix as the linear weighted matrix. Then, each column element in the above-mentioned linear weighted matrix can be added to the above-mentioned image bias vector, and the resulting matrix is used as the video frame feature information. Among them, the linear matrix can be a matrix that can be left-multiplied by the above-mentioned weighted matrix and is used to map the above-mentioned weighted matrix to a preset feature dimension. As an example, when the feature dimension to be mapped is 3D and the above-mentioned weighted matrix is a 2-row and 2-column matrix, the number of rows of the linear matrix corresponds to the above-mentioned feature dimension, and the number of columns of the linear matrix is the number of rows of the above-mentioned weighted matrix, that is, the linear matrix is a 3-row and 2-column matrix. The above-mentioned image bias vector can be an array with the same number of rows as the above-mentioned linear weighted matrix and a corresponding column number of 1.

[0058] The second layer can be a text processing layer. The above-mentioned text processing layer can include a text mapping network, a text embedding layer, and a text linear layer. Among them, the above-mentioned text mapping network can be a neural network that takes masked question text data as input and outputs the question text feature information corresponding to the masked question text data. For example, the above-mentioned text mapping network can be a recurrent neural network. The above-mentioned question text feature information can be a feature vector corresponding to the masked question text data.

[0059] The above-mentioned text embedding layer can be a fully connected layer that takes the question text feature information corresponding to the masked question text data as input and outputs the text embedding feature information corresponding to the masked question text data. The above-mentioned text embedding feature information can be the question text feature information processed by the text embedding layer. The above-mentioned text embedding layer corresponds to a pre-trained weight matrix.

[0060] In practice, first, the above-mentioned text embedding layer can fine-tune the pre-trained weight matrix through parameter fine-tuning technology to obtain the fine-tuned pre-trained weight matrix as the matrix to be processed. Then, the above-mentioned text embedding layer can determine the product of the above-mentioned matrix to be processed and the question text feature information as the text embedding feature information. Among them, the above-mentioned parameter fine-tuning technology can be a technology that can fine-tune the parameters of a neural network. For example, the above-mentioned parameter fine-tuning technology can be LoRA (Low-Rank Adaptation). The above-mentioned pre-trained weight matrix can be a parameter matrix used to process the question text feature information in the above-mentioned text embedding layer.

[0061] The above-mentioned text linear layer can be a linear layer that takes the text embedding feature information corresponding to the masked question text data as input and outputs the video text feature information corresponding to the masked question text data. Among them, the above-mentioned video text feature information can be the text embedding feature information after linear transformation. The above-mentioned text linear layer corresponds to a text weight matrix.

[0062] In practice, first, the above-mentioned text linear layer can determine the product of the text weight matrix and the above-mentioned text embedding feature information as the text linear matrix. Then, each column element in the above-mentioned text linear matrix can be added to a preset text bias vector, and the resulting matrix after addition is used as the video text feature information. Among them, the text weight matrix can be a parameter matrix that can be left-multiplied by the text embedding feature information and is used to map the text embedding feature information to the above-mentioned feature dimension. The above-mentioned text bias vector can be an array with the same number of rows as the text linear matrix and a column number of 1.

[0063] The third layer can be the output layer. Among them, the above-mentioned output layer can include a normalization activation function and an output function. The above-mentioned normalization activation function can be a Softmax function that takes the video frame feature information output by the above-mentioned image linear layer and the video text feature information output by the above-mentioned text linear layer as inputs and outputs row-wise probability distribution data. The above-mentioned behavior probability distribution data can be used to represent the probability distribution of the behavior corresponding to the target video frame. The above-mentioned behavior probability distribution data can include behavior information and behavior probability. The above-mentioned behavior information can be used to describe the behavior corresponding to the above-mentioned target video frame. For example, the above-mentioned behavior information can be "running". The above-mentioned behavior probability can be the accuracy rate corresponding to the above-mentioned behavior information. For example, the above-mentioned behavior probability distribution data can be {"behavior information: running; behavior probability: 70%", "behavior information: walking; behavior probability: 30%"}.

[0064] The above-mentioned output function can be an argmax function that takes the behavior probability distribution data output by the above-mentioned normalization activation function as an input and outputs behavior description information. Among them, the above-mentioned behavior description information can be the behavior information with the largest corresponding behavior probability in the behavior probability distribution data.

[0065] In some optional implementation manners of some embodiments, the above-mentioned behavior recognition model can be obtained by the above-mentioned execution subject through the following steps: The first step is to obtain a sample set. Among them, each sample in the sample set includes a sample video frame, sample masked question text data, and sample behavior description information. The above-mentioned sample video frame can be the target video frame for model training. The above-mentioned sample masked question text data can be the masked question text data for model training. The above-mentioned sample behavior description information can be the standard behavior description information corresponding to the sample video frame and the sample masked question text data. The above-mentioned sample behavior description information includes each sample token. Each sample token in the above-mentioned each sample token can be a token in the above-mentioned sample behavior description information.

[0066] The second step is to perform the following training steps based on the sample set: The first sub-step is to input each sample video frame included in at least one sample in the sample set into the image mapping network in the image processing layer of the initial neural network to obtain each sample video frame information. Among them, each sample video frame information in the above-mentioned each sample video frame information can be the video frame information corresponding to the sample video frame. Each sample video frame information in the above-mentioned each sample video frame information corresponds to a sample. The structure of the initial neural network can refer to the above-mentioned behavior recognition model and will not be elaborated here.

[0067] The second sub-step is to input each sample video frame information corresponding to the above-mentioned at least one sample into the image embedding network in the image processing layer of the initial neural network to obtain each sample weighted matrix. Among them, each sample weighted matrix in the above-mentioned each sample weighted matrix can be the weighted matrix corresponding to the sample video frame. Each sample weighted matrix in the above-mentioned each sample weighted matrix corresponds to a sample.

[0068] The third sub-step is to input each sample weighted matrix corresponding to the above-mentioned at least one sample into the image linear layer in the image processing layer of the initial neural network to obtain each sample video frame feature information corresponding to the above-mentioned at least one sample. Among them, each sample video frame feature information in the above-mentioned each sample video frame feature information can be the video frame feature information corresponding to the sample video frame.

[0069] The fourth sub-step is to input each sample mask problem text data included in the above-mentioned at least one sample into the text mapping network in the text processing layer of the initial neural network to obtain each sample problem text feature information. Among them, each sample problem text feature information in the above-mentioned each sample problem text feature information can be the problem text feature information corresponding to the sample video frame. Each sample problem text feature information in the above-mentioned each sample problem text feature information corresponds to a sample.

[0070] The fifth sub-step is to input each sample problem text feature information corresponding to the above-mentioned at least one sample into the text embedding layer in the text processing layer of the initial neural network to obtain each sample text embedding feature information. Among them, each sample text embedding feature information in the above-mentioned each sample text embedding feature information can be the text embedding feature information corresponding to the sample video frame. Each sample text embedding feature information in the above-mentioned each sample text embedding feature information corresponds to a sample.

[0071] Sixth sub-step: Input the feature information embedded in the text of each sample corresponding to the above at least one sample into the text linear layer in the text processing layer of the initial neural network to obtain the text feature information of each sample video. Among them, each sample video text feature information in the above-mentioned each sample video text feature information may be the video text feature information corresponding to the sample video frame. Each sample video text feature information in the above-mentioned each sample video text feature information corresponds to one sample.

[0072] Seventh sub-step: For each sample in the above at least one sample, based on the above-mentioned each sample video frame feature information and the above-mentioned each sample video text feature information, generate the alignment loss data corresponding to the above sample. Among them, the above alignment loss data may be the loss value corresponding to the above sample.

[0073] In practice, first, the feature information of each sample video frame included in the above at least one sample can be determined as the feature information of each target video frame. Secondly, for each sample in the above at least one sample, based on the sample video text feature information corresponding to the above sample, the feature similarity between the above sample video text feature information and each target video frame feature information in the above each target video frame feature information can be determined as the feature similarity, and each feature similarity is obtained. Then, for each feature similarity in the above each feature similarity, the power of the preset exponential base data corresponding to the above feature similarity can be determined as the feature exponential data. Among them, the above exponential base data may be the natural constant e. As an example, when the above feature similarity is 3, the corresponding feature exponential data is the 3rd power of e. Then, the sum of the determined each feature exponential data can be determined as the comprehensive exponential data. Then, the feature exponential data corresponding to the above sample in the above each feature exponential data can be determined as the target feature exponential data. As an example, the feature exponential data corresponding to the sample is the feature exponential data generated by the feature information of the sample video frame and the sample video text feature information included in the sample. Then, the ratio of the above target feature exponential data to the above comprehensive exponential data can be determined as the exponential ratio data. Finally, the negative logarithm of the above exponential ratio data can be determined as the alignment loss data corresponding to the above sample.

[0074] Eighth sub-step: In response to determining that the generated alignment loss data does not meet the preset loss condition, adjust the network parameters of the image processing layer and the text processing layer in the initial neural network, and use the unused samples to form a sample set. Use the image processing layer and the text processing layer of the adjusted initial neural network to execute the above training steps again.

[0075] Among them, the above preset loss condition can be that the average value of each alignment loss data is less than a preset loss value. The above preset loss value can be a preset numerical value. Here, there is no limitation on the specific setting of the above preset loss value. The network parameters corresponding to the image mapping network in the image processing layer and the text mapping network in the text processing layer are in a frozen state, that is, when adjusting the network parameters of the initial neural network, the network parameters of the image mapping network in the image processing layer and the text mapping network in the text processing layer remain unchanged.

[0076] In practice, first, the above-mentioned execution entity can determine the average value of each of the above alignment loss data as the loss value to be processed. Then, for the query weight matrix corresponding to the image embedding network in the image processing layer of the initial neural network, first, the partial derivative of the above loss value to be processed and the query weight matrix can be determined as gradient data. Then, the product of the preset learning rate, the above gradient data, and the above loss value to be processed can be determined as the query weight matrix to update the query weight matrix. For the key weight matrix and value weight matrix corresponding to the image embedding network in the image processing layer of the initial neural network, and the linear matrix in the image linear layer, the key weight matrix, value weight matrix, and linear matrix can be updated. Among them, the method of updating the key weight matrix, value weight matrix, and linear matrix can refer to the specific implementation method of updating the query weight matrix, which will not be elaborated here. Thus, the network parameters of the image processing layer in the initial neural network can be adjusted. In practice, the above-mentioned execution entity can adjust the network parameters of the text processing layer in the initial neural network through the backpropagation algorithm.

[0077] Optionally, after generating the alignment loss data corresponding to each of the above at least one sample based on the above respective sample video frame feature information and the above respective sample video text feature information for each of the above at least one sample, the above-mentioned execution entity may further perform the following steps: In response to determining that the generated respective alignment loss data meet the preset loss condition, perform the following steps: In the first step, input the respective sample video frame feature information and the respective sample video text feature information corresponding to the above at least one sample into the output layer of the initial neural network to obtain the respective behavior description information corresponding to the above at least one sample. Among them, each behavior description information in the above respective behavior description information includes respective tokens. Each of the above tokens can be a token in the behavior description information.

[0078] In the second step, for each of the at least one sample described above, based on the behavior description information corresponding to the sample in each of the above behavior description information and the sample behavior description information included in the sample, text metric data corresponding to the sample is generated. Wherein, the text metric data may be a numerical value used to characterize the accuracy of the generated behavior description information.

[0079] In the third step, in response to determining that the generated text metric data satisfies a preset metric data condition, the initial neural network is determined as a behavior recognition model. Wherein, the metric data condition may be that the average value of each of the above text metric data is greater than a preset metric value. The preset metric value may be a numerically set value in advance. Here, no limitation is imposed on the specific setting of the preset metric value.

[0080] In the fourth step, in response to determining that the generated text metric data does not satisfy the above metric data condition, the network parameters of the output layer in the initial neural network are adjusted, and a sample set is formed using the unused samples. The output layer of the adjusted initial neural network is used to perform the above training steps again. In practice, the network parameters of the output layer in the initial neural network can be adjusted by using the backpropagation algorithm.

[0081] In the process of adopting technical solutions to solve the above technical problems, the following problems often accompany: When performing behavior recognition on a video through a feature extraction algorithm and a classifier, it is easily interfered by the background information in the video, resulting in a low recognition accuracy, and further causing a waste of computing resources when it is necessary to consume computing resources again to re-recognize the recognized video.

[0082] In the face of the above technical problems, the following solution is decided to be adopted: In some optional implementation manners of some embodiments, the above execution subject may generate text metric data corresponding to the sample based on the behavior description information corresponding to the sample in each of the above behavior description information and the sample behavior description information included in the sample through the following steps: In the first step, each token included in the above behavior description information is determined as a token sequence. In practice, the above execution subject may arrange the above tokens in the order of the above tokens in the above behavior description information, from front to back, to form a token sequence.

[0083] In the second step, each sample token included in the above sample behavior description information is determined as a sample token sequence. In practice, the above execution subject may arrange the above sample tokens in the order of the above sample tokens in the above sample behavior description information, from front to back, to form a sample token sequence.

[0084] In the third step, for each of the preset token quantities, perform the following steps: The first sub-step is to generate respective token subsequences based on the above-mentioned token quantity and the above-mentioned token sequence. Among them, each token subsequence in the above-mentioned respective token subsequences can be a subsequence of the above-mentioned token sequence. In practice, for every adjacent above-mentioned token quantity of tokens in the above-mentioned token sequence, the above-mentioned execution entity can determine the above-mentioned token quantity of tokens as a token subsequence, and thus, obtain respective token subsequences.

[0085] The second sub-step is to generate respective sample token subsequences based on the above-mentioned token quantity and the above-mentioned sample token sequence. Among them, each sample token subsequence in the above-mentioned respective sample token subsequences can be a subsequence of the above-mentioned sample token sequence. In practice, for every adjacent above-mentioned token quantity of sample tokens in the above-mentioned sample token sequence, the above-mentioned token quantity of sample tokens can be determined as a sample token subsequence, and thus, respective sample token subsequences can be obtained.

[0086] The third sub-step is to, for each token subsequence in the above-mentioned respective token subsequences, perform the following steps: Sub-step one, determine the above-mentioned token subsequence as the target token subsequence.

[0087] Sub-step two, based on the above-mentioned target token subsequence and the above-mentioned respective token subsequences, determine the token quantity information corresponding to the above-mentioned target token subsequence. Among them, the above-mentioned token quantity information can be the quantity of the respective token subsequences in the above-mentioned respective token subsequences that are the same as the above-mentioned target token subsequence. In practice, the above-mentioned execution entity can generate the number of the respective token subsequences in the above-mentioned respective token subsequences that are the same as the above-mentioned target token subsequence as the token quantity information through a string function. Among them, the above-mentioned string function can be a function that can generate the number of times a substring appears in a specified string. For example, the above-mentioned string function can be str.count().

[0088] Sub-step three, based on the above-mentioned target token subsequence and the above-mentioned respective sample token subsequences, determine the sample quantity information corresponding to the above-mentioned target token subsequence. Among them, the above-mentioned sample quantity information can be the quantity of the respective sample token subsequences in the above-mentioned respective sample token subsequences that have the same content as the above-mentioned target token subsequence. In practice, the above-mentioned execution entity can generate the number of the respective sample token subsequences in the above-mentioned respective sample token subsequences that are the same as the above-mentioned target token subsequence as the sample quantity information through the above-mentioned string function.

[0089] Sub-step 4: Based on the above-mentioned token quantity information and the above-mentioned sample quantity information, determine the target quantity information corresponding to the above-mentioned target token subsequence. Among them, the above-mentioned target quantity information can be the above-mentioned token quantity information or the above-mentioned sample quantity information. In practice, in response to determining that the above-mentioned token quantity information is greater than or equal to the above-mentioned sample quantity information, the above-mentioned sample quantity information can be determined as the target quantity information. In response to determining that the above-mentioned token quantity information is less than the above-mentioned sample quantity information, the token quantity information can be determined as the target quantity information.

[0090] Fourth sub-step: Determine the sum of the determined target quantity information of each item as the sum quantity information.

[0091] Fifth sub-step: Determine the quantity corresponding to each of the above-mentioned token subsequences as the total quantity information.

[0092] Sixth sub-step: Determine the ratio of the above-mentioned sum quantity information to the above-mentioned total quantity information as the token precision data corresponding to the above-mentioned token quantity.

[0093] Seventh sub-step: Determine the quantity corresponding to each of the above-mentioned sample token subsequences as the sample total quantity information.

[0094] Eighth sub-step: Determine the ratio of the above-mentioned sum quantity information to the above-mentioned sample total quantity information as the token recall data corresponding to the above-mentioned token quantity.

[0095] Fourth step: Determine the token quantity that meets the preset quantity condition among the above-mentioned token quantities as the target token quantity. Among them, the above-mentioned preset quantity condition can be that the value of the token quantity is the largest among the above-mentioned token quantities.

[0096] Fifth step: Determine the reciprocal of the above-mentioned target token quantity as the token weight data.

[0097] Sixth step: Based on each token precision data among the determined token precision data, determine the logarithm of the above-mentioned token precision data as the token logarithm data.

[0098] Seventh step: Based on each token logarithm data among the determined token logarithm data, determine the product of the above-mentioned token logarithm data and the above-mentioned token weight data as the token product data.

[0099] Eighth step: Determine the sum of the determined token product data of each item as the token total data.

[0100] Step 9: Generate token exponential data based on the preset exponential base data and the above-mentioned total token data. Among them, the above exponential base data can be the natural constant e. The above token exponential data can be a value generated from the above exponential base data and the above total token data. In practice, the above execution entity can determine the power of the above exponential base data to the above total token data as the token exponential data. For example, when the above total token data is 3, the corresponding token exponential data is the 3rd power of e.

[0101] Step 10: Generate token difference data based on the above behavior description information, the above sample behavior description information, and preset data. Among them, the above token difference data can be a value generated from the above behavior description information, the above sample behavior description information, and the above preset data. The above preset data can be 1. In practice, first, the above execution entity can determine the number of each token included in the above behavior description information as the token number data. Second, the number of each sample token included in the above sample behavior description information can be determined as the sample token number data. Then, in response to determining that the above token number data is greater than or equal to the above sample token number data, the above sample token number data can be determined as the target token number data. In response to determining that the above token number data is less than the above sample token number data, the above token number data can be determined as the target token number data. Then, the ratio of the determined target token number data to the above token number data can be determined as the token ratio data. Finally, the difference between the above preset data and the above token ratio data can be determined as the token difference data.

[0102] Step 11: Generate a weight factor data based on the above exponential base data and the above token difference data. Among them, the above weight factor data can be a value generated from the above exponential base data and the above token difference data. In practice, the above execution entity can determine the power of the above exponential base data to the above token difference data as the weight factor data. For example, when the above exponential base data is e and the above token difference data is 4, the corresponding weight factor data is the 4th power of e.

[0103] Step 12: Determine the product of the above weight factor data and the above token exponential data as the precision index data.

[0104] Step 13: Generate recall metric data based on the retrieved data for each identified token. Among them, the above recall metric data can be the weighted average corresponding to the retrieved data for each of the above tokens. In practice, for each token retrieval data among the above token retrieval data, first, the above execution entity can determine the reciprocal of the number of tokens corresponding to the above token retrieval data as the token reciprocal data. Then, the product of the above token reciprocal data and the above token retrieval data can be determined as the retrieval data. Finally, the sum of the obtained retrieval data can be determined as the recall metric data.

[0105] Step 14: Generate text accuracy data based on the above behavior description information and the above sample behavior description information. Among them, the above text accuracy data can be a value used to characterize the accuracy of the above behavior description information compared to the above sample behavior description information. In practice, for each token included in the above behavior description information, the above execution entity can retrieve the above token from the above sample behavior description information through the above string function to obtain the return value output by the above string function. Among them, the above return value can be a null value or the number of occurrences. The above number of occurrences can be the number of times the above token appears in the above sample behavior description information. Then, each of the return values that meet the preset return condition among the obtained return values can be determined as each target return value. Among them, the above preset return condition can be that the return value is not a null value. Then, the number corresponding to each of the above target return values can be determined as the return value quantity data. Then, the number of each token included in the above behavior description information can be determined as the to-be-processed quantity data. Finally, the ratio of the above return value quantity data to the above to-be-processed quantity data can be determined as the text accuracy data.

[0106] Step 15: Generate text metric data based on the above accuracy metric data, the above recall metric data, and the above text accuracy data. In practice, the above execution entity can determine the average value of the above accuracy metric data, the above recall metric data, and the above text accuracy data as the text metric data.

[0107] The above technical solution and its related content in combination with steps 101 to 109 are an inventive point of an embodiment of the present disclosure, which solves the problem of "waste of computing resources". The factors that lead to relatively large consumption of computing resources are often as follows: When performing behavior recognition on a video through a feature extraction algorithm and a classifier, it is easily interfered by the background information in the video, resulting in a low recognition accuracy. As a result, when it is necessary to re-consume computing resources to re-recognize the recognized video, there is a waste of computing resources. If the above factors are solved, the waste of computing resources can be reduced. To achieve this effect, the present disclosure first determines each token included in the above behavior description information as a token sequence. Secondly, each sample token included in the above sample behavior description information is determined as a sample token sequence. Thus, the data to be processed can be obtained. Then, for each token number among the preset token numbers, the following steps are performed: First, based on the above token number and the above token sequence, each token subsequence is generated. Thus, each token subsequence corresponding to the above token number can be obtained. Secondly, based on the above token number and the above sample token sequence, each sample token subsequence is generated. Thus, each sample token subsequence corresponding to the above token number can be obtained. Then, for each token subsequence among the above token subsequences, the following steps are performed: First, the above token subsequence is determined as the target token subsequence. Thus, the target token subsequence can be obtained. Secondly, based on the above target token subsequence and the above token subsequences, the token number information corresponding to the above target token subsequence is determined. Thus, the number of token subsequences that are the same as the target token subsequence among the above token subsequences can be obtained. Then, based on the above target token subsequence and the above sample token subsequences, the sample number information corresponding to the above target token subsequence is determined. Thus, the number of sample token subsequences that are the same as the target token subsequence among the above sample token subsequences can be obtained. Then, based on the above token number information and the above sample number information, the target number information corresponding to the above target token subsequence is determined. Then, the sum of the determined target number information is determined as the sum number information. Thus, the sum number information can be obtained. Then, the number corresponding to each of the above token subsequences is determined as the total number information. Thus, the total number information can be obtained. Then, the ratio of the above sum number information to the above total number information is determined as the token precision data corresponding to the above token number. Thus, the precision of the generated behavior description information can be obtained. Then, the number corresponding to each of the above sample token subsequences is determined as the sample total number information. Thus, the sample total number information can be obtained. Then, the ratio of the above sum number information to the above sample total number information is determined as the token recall data corresponding to the above token number. Thus, the recall rate of the generated behavior description information can be obtained. Then, the token number that satisfies the preset number condition among the above token numbers is determined as the target token number.Then, determine the reciprocal of the above-mentioned number of target tokens as the token weight data. Then, based on each token accuracy data among the determined token accuracy data, determine the logarithm of the above-mentioned token accuracy data as the token logarithm data. Then, based on each token logarithm data among the determined token logarithm data, determine the product of the above-mentioned token logarithm data and the above-mentioned token weight data as the token product data. Then, determine the sum of the determined token product data as the token sum data. Thus, the token sum data can be obtained. Then, based on the preset exponential base data and the above-mentioned token sum data, generate the token exponential data. Thus, the token exponential data can be obtained. Then, based on the above-mentioned behavior description information, the above-mentioned sample behavior description information, and the preset data, generate the token difference data. Then, based on the above-mentioned exponential base data and the above-mentioned token difference data, generate the weight factor data. Thus, the weight factor data can be obtained. Then, determine the product of the above-mentioned weight factor data and the above-mentioned token exponential data as the accuracy index data. Thus, the accuracy index data can be obtained. Then, based on the determined token recall data, generate the recall index data. Thus, the weighted average of the respective token recall data can be obtained. Then, based on the above-mentioned behavior description information and the above-mentioned sample behavior description information, generate the text accuracy data. Thus, the accuracy rate of the generated behavior description information can be obtained. Finally, based on the above-mentioned accuracy index data, the above-mentioned recall index data, and the above-mentioned text accuracy data, generate the text index data. Thus, the corresponding quality index of the generated behavior description information can be obtained. Also, because the network parameters of the initial neural network can be adjusted according to the quality index of the generated behavior description information, and the accuracy rate of the behavior description information generated by the trained behavior recognition model can be improved through continuous iteration, the probability of the situation where it is necessary to consume computing resources to perform behavior recognition on the recognized video again due to a low accuracy rate can be reduced, and thus the waste of computing resources can be reduced.

[0108] Step 107: Input the masked problem text data into the text processing layer in the behavior recognition model to obtain video text feature information.

[0109] In some embodiments, the above-mentioned execution subject may input the above-mentioned masked problem text data into the text processing layer in the above-mentioned behavior recognition model to obtain video text feature information. Among them, the above-mentioned text processing layer includes a text mapping network, a text embedding layer, and a text linear layer.

[0110] Step 108: Input the video frame feature information and the video text feature information into the output layer in the behavior recognition model to obtain behavior description information.

[0111] In some embodiments, the execution entity may input the video frame feature information and the video text feature information into an output layer of the behavior recognition model to obtain behavior description information.

[0112] Step 109: Send the behavior description information to a user terminal.

[0113] In some embodiments, the execution entity may send the behavior description information to the user terminal.

[0114] The above embodiments of the present disclosure have the following beneficial effects: Through the large-scale video behavior recognition method based on Fourier transform significance in some embodiments of the present disclosure, waste of computing resources can be reduced. Specifically, the reason for the waste of computing resources is that when recognizing behaviors in a video, the recognition process is easily interfered by the background information in the video, resulting in a low recognition accuracy. As a result, the probability of consuming computing resources to recognize the video again is relatively high, leading to waste of computing resources. Moreover, when using a feature extraction algorithm to extract features from a video, each frame in the video needs to be processed, resulting in a large amount of computing resources consumed during processing. Based on this, in some embodiments of the large-scale video behavior recognition method based on Fourier transform significance of the present disclosure, first, in response to receiving the video data to be processed and the problem text data sent by the user terminal, a group of video frames to be processed corresponding to the video data to be processed is generated. Thus, the video data to be recognized can be obtained. Secondly, based on the group of video frames to be processed, various frequency domain transformation matrices corresponding to the group of video frames to be processed are generated. Thus, by performing Fourier transform processing on the video, the model can better recognize foreground information during recognition and reduce interference from background information. Then, based on the various frequency domain transformation matrices, the target video frame corresponding to the video data to be processed is determined. Thus, the video frame to be processed with the highest significance can be determined from the group of video frames to be processed. Then, based on the problem text data, target problem text data is generated. Then, mask processing is performed on the target problem text data to obtain masked problem text data. Thus, redundant information in the text data can be masked. Then, the target video frame is input into the image processing layer of the pre-trained behavior recognition model to obtain video frame feature information, where the image processing layer includes an image mapping network, an image embedding network, and an image linear layer. Thus, the target video frame can be processed by the behavior recognition model. Then, the masked problem text data is input into the text processing layer of the behavior recognition model to obtain video text feature information, where the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer. Thus, the masked problem text data can be processed by the behavior recognition model. Then, the video frame feature information and the video text feature information are input into the output layer of the behavior recognition model to obtain behavior description information. Thus, the behavior description information corresponding to the video can be obtained. Finally, the behavior description information is sent to the user terminal. Also, because a most significant video frame to be processed can be determined first from the group of video frames to be processed of the video data to be processed, and then the most significant video frame is processed to obtain the behavior description information, instead of processing each video frame to be processed in the video data to be processed, the computing resources consumed during processing can be reduced.It is also because the video can be processed by Fourier transform first, enabling the model to better process foreground information and reducing the interference of background information on recognition. Therefore, the recognition accuracy can be improved, and further, the probability of the situation where computational resources need to be consumed to recognize the video again can be reduced, and thus the waste of computational resources can be reduced.

[0115] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a large-scale video behavior recognition device based on Fourier transform saliency. These device embodiments correspond to Figure 2 the method embodiments shown, and the device can be specifically applied to various electronic devices.

[0116] As Figure 2 shown, some embodiments of the large-scale video behavior recognition device 200 based on Fourier transform saliency include: a first generation unit 201, a second generation unit 202, a determination unit 203, a third generation unit 204, a masking unit 205, a first input unit 206, a second input unit 207, a third input unit 208, and a sending unit 209. Among them, the first generation unit 201 is configured to generate a group of video frames to be processed corresponding to the video data to be processed in response to receiving the video data to be processed and the question text data sent by the user terminal; the second generation unit 202 is configured to generate respective frequency domain transformation matrices corresponding to the group of video frames to be processed based on the group of video frames to be processed; the determination unit 203 is configured to determine a target video frame corresponding to the video data to be processed based on the respective frequency domain transformation matrices; the third generation unit 204 is configured to generate target question text data based on the question text data; the masking unit 205 is configured to perform masking processing on the target question text data to obtain masked question text data; the first input unit 206 is configured to input the target video frame into an image processing layer in a pre-trained behavior recognition model to obtain video frame feature information, where the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; the second input unit 207 is configured to input the masked question text data into a text processing layer in the behavior recognition model to obtain video text feature information, where the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; the third input unit 208 is configured to input the video frame feature information and the video text feature information into an output layer in the behavior recognition model to obtain behavior description information; the sending unit 209 is configured to send the behavior description information to the user terminal.

[0117] It can be understood that the various units described in the large-scale video behavior recognition device 200 based on Fourier transform saliency and the reference Figure 1corresponds to each step in the described method. Thus, the operations, features, and beneficial effects described above for the method also apply to the large-scale video behavior recognition device 200 based on Fourier transform saliency and the units included therein, which will not be elaborated here.

[0118] Reference is now made to Figure 3 , which shows a schematic structural diagram of an electronic device (such as a computing device) 300 suitable for use in implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.

[0119] As Figure 3 shown, the electronic device 300 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0120] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wirelessly to exchange data. Although Figure 3 shows an electronic device 300 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included. Figure 3 Each block shown in

[0121] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product that includes a computer program carried on a computer-readable medium, and the computer program contains program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program can be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are performed.

[0122] It should be noted that the computer-readable medium described in some embodiments of the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0123] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0124] The computer-readable medium described above can be included in the above-mentioned electronic device; it can also exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to receiving the video data to be processed and the problem text data sent by the user terminal, generate a group of video frames to be processed corresponding to the video data to be processed; based on the group of video frames to be processed, generate respective frequency-domain transformation matrices corresponding to the group of video frames to be processed; based on the respective frequency-domain transformation matrices, determine the target video frame corresponding to the video data to be processed; based on the problem text data, generate the target problem text data; perform masking processing on the target problem text data to obtain the masked problem text data; input the target video frame into the image processing layer of a pre-trained behavior recognition model to obtain video frame feature information, where the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; input the masked problem text data into the text processing layer of the behavior recognition model to obtain video text feature information, where the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; input the video frame feature information and the video text feature information into the output layer of the behavior recognition model to obtain behavior description information; and send the behavior description information to the user terminal.

[0125] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, execute as a stand-alone software package, execute partially on the user's computer and partially on a remote computer, or execute entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0126] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0127] The units described in some embodiments of the present disclosure may be implemented in software or in hardware. The described units may also be provided in a processor. For example, a processor may be described as including a first generation unit, a second generation unit, a determination unit, a third generation unit, a masking unit, a first input unit, a second input unit, a third input unit, and a sending unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the sending unit may also be described as "the unit that sends the above-described behavior description information to the above user terminal".

[0128] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), System on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.

[0129] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. A large-scale video behavior recognition method based on Fourier transform saliency, comprising: Responding to receiving the to-be-processed video data and problem text data sent by the user terminal, generating a to-be-processed video frame group corresponding to the to-be-processed video data; Based on the to-be-processed video frame group, generating respective frequency domain transformation matrices corresponding to the to-be-processed video frame group; Based on the respective frequency domain transformation matrices, determining a target video frame corresponding to the to-be-processed video data; Based on the problem text data, generating target problem text data; Performing mask processing on the target problem text data to obtain masked problem text data; Inputting the target video frame into an image processing layer in a pre-trained behavior recognition model to obtain video frame feature information, wherein the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; Inputting the masked problem text data into a text processing layer in the behavior recognition model to obtain video text feature information, wherein the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; Inputting the video frame feature information and the video text feature information into an output layer in the behavior recognition model to obtain behavior description information; Sending the behavior description information to the user terminal.

2. The method according to claim 1, wherein, Each to-be-processed video frame in the to-be-processed video frame group corresponds to image height data and image width data; And the generating respective frequency domain transformation matrices corresponding to the to-be-processed video frame group based on the to-be-processed video frame group includes: For each to-be-processed video frame in the to-be-processed video frame group, performing the following steps: Based on a preset first gray coefficient value, a preset second gray coefficient value, and a preset third gray coefficient value, performing image conversion processing on the to-be-processed video frame to obtain gray video frame data; Performing frequency domain transformation processing on the gray video frame data to obtain a frequency domain transformation matrix.

3. The method according to claim 2, wherein, The performing frequency domain transformation processing on the gray video frame data to obtain a frequency domain transformation matrix includes: Based on the image height data and image width data corresponding to the to-be-processed video frame, determining respective matrix height data and respective matrix width data; Based on the gray video frame data, determining a gray video matrix; For each matrix height data in the respective matrix height data and each matrix width data in the respective matrix width data, performing the following steps: For each element included in the gray video matrix, performing the following steps: Based on the gray video matrix and the element, determining element height data and element width data corresponding to the element; Determining matrix element height data as the product of the matrix height data and the element height data; Determining matrix element width data as the product of the matrix width data and the element width data; Determining height ratio data as the ratio of the matrix element height data to the image height data corresponding to the to-be-processed video frame; Determining width ratio data as the ratio of the matrix element width data to the image width data corresponding to the to-be-processed video frame; Determine the sum of the height ratio data and the width ratio data as the sum ratio data; Determine the product of the preset imaginary data, the preset coefficient data, and the sum ratio data as the exponential data; Generate the to-be-processed exponential data based on the preset exponential base data and the exponential data; Determine the product of the element value corresponding to the element and the to-be-processed exponential data as the matrix sub-element data; Determine the sum of the determined matrix sub-element data as the matrix element data corresponding to the matrix height data and the matrix width data; Combine the obtained matrix element data into a frequency domain transformation matrix.

4. The method according to claim 1, wherein, The determining the target video frame corresponding to the to-be-processed video data based on the respective frequency domain transformation matrices includes: For each frequency domain transformation matrix in the respective frequency domain transformation matrices, perform the following steps: Generate a spectral energy matrix based on the frequency domain transformation matrix; Generate spectral element accumulation data based on the spectral energy matrix; Determine the product of the image height data and the image width data corresponding to the to-be-processed video frame as the image processing data; Determine the ratio of the spectral element accumulation data to the image processing data as the average energy data; Determine the average energy data that satisfies the preset energy data condition among the determined average energy data as the target average energy data; Determine the to-be-processed video frame corresponding to the target average energy data in the to-be-processed video frame group as the target video frame.

5. The method according to claim 1, wherein, The behavior recognition model is trained through the following steps: Obtain a sample set, where each sample in the sample set includes a sample video frame, sample mask question text data, and sample behavior description information, and each sample word element is included in the sample behavior description information; Based on the sample set, perform the following training steps: Input each sample video frame included in at least one sample in the sample set into the image mapping network in the image processing layer of the initial neural network to obtain each sample video frame information, where each sample video frame information in the each sample video frame information corresponds to a sample; Input the each sample video frame information corresponding to the at least one sample into the image embedding network in the image processing layer of the initial neural network to obtain each sample weighted matrix, where each sample weighted matrix in the each sample weighted matrix corresponds to a sample; Input the each sample weighted matrix corresponding to the at least one sample into the image linear layer in the image processing layer of the initial neural network to obtain the each sample video frame feature information corresponding to the at least one sample; Input the each sample mask question text data included in the at least one sample into the text mapping network in the text processing layer of the initial neural network to obtain each sample question text feature information, where each sample question text feature information in the each sample question text feature information corresponds to a sample; Input the text feature information of each sample problem corresponding to the at least one sample into the text embedding layer in the text processing layer of the initial neural network to obtain the text embedding feature information of each sample, where each text embedding feature information in the text embedding feature information of each sample corresponds to one sample; Input the text embedding feature information of each sample corresponding to the at least one sample into the text linear layer in the text processing layer of the initial neural network to obtain the video text feature information of each sample, where each video text feature information in the video text feature information of each sample corresponds to one sample; For each sample in the at least one sample, generate alignment loss data corresponding to the sample based on the video frame feature information of each sample and the video text feature information of each sample; In response to determining that the generated alignment loss data does not meet the preset loss condition, adjust the network parameters of the image processing layer and the text processing layer in the initial neural network, and use the unused samples to form a sample set, and use the adjusted image processing layer and text processing layer of the initial neural network to execute the training step again.

6. The method according to claim 5, wherein After generating the alignment loss data corresponding to each sample for each sample in the at least one sample based on the video frame feature information of each sample and the video text feature information of each sample, the method further includes: In response to determining that the generated alignment loss data meets the preset loss condition, execute the following steps: Input the video frame feature information of each sample and the video text feature information of each sample corresponding to the at least one sample into the output layer in the initial neural network to obtain the behavior description information of each sample corresponding to the at least one sample, where each behavior description information includes each token; For each sample in the at least one sample, generate text metric data corresponding to the sample based on the behavior description information corresponding to the sample in the behavior description information of each sample and the sample behavior description information included in the sample; In response to determining that the generated text metric data meets the preset metric data condition, determine the initial neural network as a behavior recognition model; In response to determining that the generated text metric data does not meet the metric data condition, adjust the network parameters of the output layer in the initial neural network, and use the unused samples to form a sample set, and use the adjusted output layer of the initial neural network to execute the training step again.

7. A large-scale video behavior recognition device based on Fourier transform saliency, comprising: A first generating unit configured to generate a group of to-be-processed video frames corresponding to the to-be-processed video data in response to receiving the to-be-processed video data and problem text data sent by a user terminal; A second generating unit configured to generate respective frequency domain transformation matrices corresponding to the group of to-be-processed video frames based on the group of to-be-processed video frames; A determining unit configured to determine a target video frame corresponding to the to-be-processed video data based on the respective frequency domain transformation matrices; A third generation unit, configured to generate target problem text data based on the problem text data; A masking unit, configured to perform masking processing on the target problem text data to obtain masked problem text data; A first input unit, configured to input the target video frame into an image processing layer in a pre-trained behavior recognition model to obtain video frame feature information, where the image processing layer includes an image mapping network, an image embedding network, and an image linear layer; A second input unit, configured to input the masked problem text data into a text processing layer in the behavior recognition model to obtain video text feature information, where the text processing layer includes a text mapping network, a text embedding layer, and a text linear layer; A third input unit, configured to input the video frame feature information and the video text feature information into an output layer in the behavior recognition model to obtain behavior description information; A sending unit, configured to send the behavior description information to the user terminal.

8. An electronic device, comprising: One or more processors; A storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable medium having a computer program stored thereon, wherein, The program, when executed by the processor, implements the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Video saliency target detection method based on frequency domain prior

    CN111178188A

  • Data processing method and device, equipment and medium

    CN116246213A

  • Video behavior recognition method, device and equipment based on multi-mode large model fine tuning

    CN119495127A