Video Key Point Recognition Method and Device, Storage Medium and Electronic Device
The video recognition model obtains and updates the face key point data in the video, and combines the optical flow data, the problems of low accuracy and high complexity of video key point recognition in the existing technology are solved, and more efficient and accurate video key point recognition is achieved.
Patent Information
- Application Number
- CN202210242640.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-11
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-03-11
AI Technical Summary
In the prior art, the face tracking method in videos is poor in accuracy and high in complexity, making it difficult to effectively identify key points in videos.
The video recognition model obtains the initial key point data and optical flow data of the multi-frame image, and updates the video recognition model with the initial and target key point data until the target key point data meets the preset conditions.
The complexity of video key points recognition is reduced, the recognition accuracy is improved, and the accuracy of the recognition process is improved by continuously updating the video recognition model.
Smart Images

Figure CN114612829B_ABST
Abstract
Description
Background Art
[0002] With the continuous development of computer science, the application scope of face recognition technology is also getting wider and wider. When tracking and detecting faces in video images, it mainly depends on the determination of face key points.
[0003] The face tracking in the prior art video is jointly realized based on a face recognition model and an optical flow estimation model. However, the method in the prior art has poor accuracy and high complexity.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a video key point recognition method, a video key point recognition device, a computer-readable medium, and an electronic device, so as to at least to some extent reduce the complexity of video key point detection and improve the recognition accuracy.
[0006] According to the first aspect of the present disclosure, there is provided a video key point recognition method, including: obtaining initial key point data of multiple frames of images in the video and optical flow data between multiple frames of images by using a video recognition model; determining target key point data corresponding to each frame of image by using the initial key point data and the optical flow data; updating the video recognition model according to the initial key point data and the target key point data; and outputting the target key point data in response to the target key point data meeting a preset condition.
[0007] According to the second aspect of the present disclosure, there is provided a video key point recognition device, including: an obtaining module, configured to obtain initial key point data of multiple frames of images in the video and optical flow data between multiple frames of images by using a video recognition model; a calculation module, configured to determine target key point data corresponding to each frame of image by using the initial key point data and the optical flow data; an updating module, configured to update the video recognition model according to the initial key point data and the target key point data; and an output module, configured to output the target key point data in response to the target key point data meeting a preset condition.
[0008] According to the third aspect of the present disclosure, there is provided a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is implemented.
[0009] According to a fourth aspect of the present disclosure, there is provided an electronic device, characterized by comprising: one or more processors; and a memory for storing one or more programs, which, when executed by the one or more processors, cause the one or more processors to implement the above-described method.
[0010] The video key point recognition method provided by an embodiment of the present disclosure uses a video recognition model to obtain initial key point data of multiple frames of images in the video and optical flow data between multiple frames of images; determines target key point data corresponding to each frame of image by using the initial key point data and the optical flow data; updates the video recognition model according to the initial key point data and the target key point data; and outputs the target key point data in response to the target key point data meeting a preset condition. Compared with the prior art, the present application updates the video recognition model according to the initial key point data and the target key point data, closely combines the optical flow data with the key points, does not require subsequent processing, reduces the complexity of video key point recognition, and at the same time, outputs the target key point data in response to the target key point data meeting the preset condition, and continuously updates the video recognition model during the recognition process until it meets the preset condition, which can improve the recognition accuracy of video key point recognition.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:
[0013] Figure 1 A schematic diagram showing an exemplary system architecture to which embodiments of the present disclosure can be applied;
[0014] Figure 2 A flowchart schematically showing a video key point recognition method in an exemplary embodiment of the present disclosure;
[0015] Figure 3 A flowchart schematically showing a process of calculating target key point data in an exemplary embodiment of the present disclosure;
[0016] Figure 4 A flowchart schematically showing a process of updating the video recognition model in an exemplary embodiment of the present disclosure;
[0017] Figure 5 Schematically shows a data flow diagram of a video key point recognition method in an exemplary embodiment of the present disclosure;
[0018] Figure 6 Schematically shows a flowchart of another video key point recognition method in an exemplary embodiment of the present disclosure;
[0019] Figure 7 Schematically shows a composition schematic diagram of a video key point recognition device in an exemplary embodiment of the present disclosure;
[0020] Figure 8 Shows a schematic diagram of an electronic device to which the embodiments of the present disclosure can be applied. Detailed implementation manners
[0021] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the example embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments.
[0022] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.
[0023] Figure 1 Shows a schematic diagram of a system architecture. The system architecture 100 can include a terminal 110 and a server 120. Among them, the terminal 110 can be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc. The server 120 generally refers to a background system that provides services related to video key point recognition in this exemplary embodiment and can be a single server or a cluster formed by multiple servers. A connection can be formed between the terminal 110 and the server 120 through a wired or wireless communication link for data interaction.
[0024] In one embodiment, the above video key point recognition method can be executed by the terminal 110. For example, after the user uses the terminal 110 to shoot a video or selects a video from the album of the terminal 110, the terminal 110 performs key point recognition on the video and outputs target key point data.
[0025] In one implementation, the above video key point recognition method can be executed by server 120. For example, after a user uses terminal 110 to shoot a video or selects a video from the album of terminal 110, terminal 110 uploads the video to server 120, and server 120 performs key point recognition on the video and returns target key point data to terminal 110.
[0026] As can be seen from the above, the execution entity of the video key point recognition method in this exemplary implementation can be the above terminal 110 or server 120, and the present disclosure does not make any limitations thereto.
[0027] The exemplary implementation of the present disclosure further provides an electronic device for executing the above video key point recognition method, and this electronic device can be the above terminal 110 or server 120. Generally, this electronic device may include a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the above image and video key point recognition method by executing the executable instructions.
[0028] Next, in conjunction with Figure 2 the video key point recognition method in this exemplary implementation will be described. Figure 2 The exemplary process of this video key point recognition method is shown and may include:
[0029] Step S210, obtaining initial key point data of multiple frames of images in the video and optical flow data between multiple frames of images by using a video recognition model;
[0030] Step S220, determining target key point data corresponding to each frame of image by using the initial key point data and the optical flow data;
[0031] Step S230, updating the video recognition model according to the initial key point data and the target key point data;
[0032] Step S240, in response to the target key point data satisfying a preset condition, outputting the target key point data.
[0033] Based on the above method, compared with the prior art, the present application updates the video recognition model according to the initial key point data and the target key point data, closely combines the optical flow data with the key points, does not require subsequent processing, reduces the complexity of video key point recognition. At the same time, in response to the target key point data satisfying a preset condition, the target key point data is output, and the video recognition model is continuously updated during the recognition process until it satisfies the preset condition, which can improve the recognition accuracy of video key point recognition.
[0034] Next, each step in Figure 2 will be specifically described.
[0035] Reference Figure 2 , in step S210, the initial key point data of multiple frames of images and the optical flow data between multiple frames of images in the video are obtained by using the video recognition model.
[0036] In an exemplary embodiment of the present disclosure, the above video recognition model may include a key point recognition sub-model and an optical flow calculation sub-model. Among them, the key point recognition sub-model is used to obtain the initial key point data in each frame of the video; the optical flow calculation sub-model is used to obtain the optical flow data between multiple frames of the video.
[0037] It should be noted that the video recognition model can recognize face key points or other key points, such as human bones, animal forms, plants, etc., and can also be customized according to user needs, which is not specifically limited in this exemplary embodiment.
[0038] In the exemplary embodiment, the above video recognition model can be obtained first, that is, the key point recognition sub-model and the optical flow calculation sub-model are obtained first.
[0039] When obtaining the above key point recognition sub-model, a first initial model can be obtained first, and then first training data can be obtained. The first training data may include multiple images and the key point data corresponding to each image. Taking the face key point data as an example, the number of key points in the face key point data can be customized according to user needs. For example, there are five points including two eyes, a nose, and two mouth corners, or it can also include more than two hundred key points, which is not specifically limited in this exemplary embodiment. The above first training data may also include video data and the key point data corresponding to each frame of the video data. The specific details of the key point data have been described in detail above, so they will not be repeated here.
[0040] The above key point recognition sub-model is mainly a neural network model based on deep learning. For example, the key point recognition sub-model can be based on a feed-forward neural network. A feed-forward network can be implemented as an acyclic graph, where nodes are arranged in layers. Generally, a feed-forward network topology includes an input layer and an output layer, and the input layer and the output layer are separated by at least one hidden layer. The hidden layer transforms the input received by the input layer into a representation useful for generating an output in the output layer. Network nodes are fully connected via edges to nodes in adjacent layers, but there are no edges between nodes within each layer. The data received at the nodes of the input layer of the feed-forward network is propagated (i.e., "fed forward") to the nodes of the output layer via an activation function, which calculates the state of the nodes in each successive layer of the network based on coefficients ("weights"), and the coefficients are respectively associated with each of the edges connecting these layers. The key point recognition sub-model can be a convolutional neural network (CNN) model or a recurrent neural network (RNN) model, but is not limited thereto, and other neural network models well-known to those skilled in the art can also be adopted.
[0041] The key point recognition sub-model needs to be obtained by training the first initial model with the above first training data. Taking the above first initial model as a convolutional neural network (CNN) model as an example, the training of the first initial model will be described below. During the supervised learning training process for the neural network, the input representing an instance in the first training data is input into the above convolutional neural network (CNN) model, and the output of the convolutional neural network (CNN) model is compared with the "correct" labeled output of this instance; an error signal representing the difference between the output and the labeled output is calculated; and when the error signal is propagated backward through the layers of the network, the weights associated with the connections are adjusted to minimize this error. When the error of each output generated from the instances in the training dataset is minimized, the first initial model is considered to be "trained" and defined as the key point recognition sub-model.
[0042] In the present exemplary embodiment, the above optical flow calculation sub-model can be a convolutional neural network (CNN) model or a recurrent neural network (RNN) model, but is not limited thereto, and other neural network models well-known to those skilled in the art can also be adopted.
[0043] Similarly, the training process of the above optical flow calculation sub-model can also adopt the above method. Specifically, first, a second initial model and corresponding second training data are obtained, where the second training data can include optical flow data between multiple frames of images in a video, and the above second training data is used to train the above second initial model to obtain the above optical flow calculation sub-model. The training process can refer to the training of the first initial model above and will not be elaborated here.
[0044] In the present exemplary embodiment, after obtaining the above-mentioned key point recognition model, that is, after obtaining the above-mentioned optical flow calculation sub-model and key point recognition sub-model, the above-mentioned video can be input into the above-mentioned optical flow calculation sub-model and key point recognition sub-model to obtain the optical flow data between multiple frames of images in the video and the initial key point data corresponding to the multiple frames of images.
[0045] In the present exemplary embodiment, the above-mentioned multiple frames of images can be the images of some frames in the video, for example, 10 frames, 20 frames, etc., or the images of all frames in the above-mentioned video, and can also be customized according to user needs, and no specific limitation is made in the present exemplary embodiment.
[0046] In step S220, the target key point data corresponding to each frame of image is determined by using the initial key point data and the optical flow data.
[0047] In an exemplary embodiment of the present disclosure, it can be set that the above-mentioned video includes M frames of images, where M is a positive integer. Referring to Figure 3 As shown, when determining the target key point data corresponding to each frame of image by using the initial key point data and the optical flow data, the following concentrated situations can be divided. Specifically, when N is equal to 1, the initial key point data of the first frame is used as the target key point data of the first frame. When N is greater than 1 and less than M, steps S310 and S320 can be included. Specifically,
[0048] In step S310, in response to N being greater than 1 and less than M, the initial key point data of the (N - a)-th frame, the initial key point data of the N-th frame, and the initial key point data of the (N + a)-th frame are obtained.
[0049] In the present exemplary embodiment, when N is greater than 1 and less than M, the initial key point data of the (N - a)-th frame, the initial key point data of the N-th frame, and the initial key point data of the (N + a)-th frame can be obtained first, where a is a positive integer, such as 1, 2, 3, etc., and can also be customized according to user needs, and no specific limitation is made in the present exemplary embodiment.
[0050] In the present exemplary embodiment, N - a is greater than or equal to 1, and N + a is less than or equal to M. Both N and M are positive integers, and M is greater than or equal to N.
[0051] In the present exemplary embodiment, the above-mentioned key point recognition sub-model can be used to obtain the initial key point data of the (N - a)-th frame, the initial key point data of the N-th frame, and the initial key point data of the (N + a)-th frame.
[0052] In the present exemplary embodiment, when a is equal to 1, relatively accurate target key point data can be calculated.
[0053] In step S320, the target key point data of the Nth frame is calculated based on at least one of the initial key point data of the (N - a)th frame and the initial key point data of the (N + a)th frame, the initial key point data of the Nth frame, and the optical flow data.
[0054] In an exemplary embodiment of the present disclosure, the target key point data can be determined based on the initial key point data of the (N - a)th frame, the initial key point data of the Nth frame, and the optical flow data. Specifically, the first intermediate key point data of the Nth frame can be calculated based on the initial key point data of the (N - a)th frame and the forward optical flow data from the (N - a)th frame to the Nth frame. Then, the average value of the first intermediate key point data and the initial key point data of the Nth frame can be used as the above-mentioned target face key data. Alternatively, the first intermediate key point data of the initial face key of the (N - a)th frame and the initial key point data of the Nth frame can be weighted and averaged to calculate the above-mentioned target key point data. The specific calculation method can also be customized according to user requirements and is not specifically limited in this exemplary embodiment.
[0055] In another exemplary embodiment of the present disclosure, the target key point data can be determined based on the initial key point data of the (N + a)th frame, the initial key point data of the Nth frame, and the optical flow data. Specifically, first, the second intermediate key point data of the Nth frame can be calculated based on the initial key point data of the (N + a)th frame and the backward optical flow data from the (N + a)th frame to the Nth frame. Then, the average value of the second intermediate key point data and the initial key point data of the Nth frame can be used as the above-mentioned target face key data. Alternatively, the second intermediate key point data and the initial key point data of the Nth frame can be weighted and averaged to calculate the above-mentioned target key point data. The specific calculation method can also be customized according to user requirements and is not specifically limited in this exemplary embodiment.
[0056] In still another exemplary embodiment of the present disclosure, the above-mentioned target key point data can be calculated based on the initial key point data of the (N - a)th frame, the initial key point data of the (N + a)th frame, the initial key point data of the Nth frame, and the optical flow data. Specifically, first, the first intermediate key point data of the Nth frame can be calculated based on the initial key point data of the (N - a)th frame and the forward optical flow data from the (N - a)th frame to the Nth frame, and the second intermediate key point data of the Nth frame can be calculated based on the initial key point data of the (N + a)th frame and the backward optical flow data from the (N + a)th frame to the Nth frame. Then, the average value of the first intermediate key point data, the second intermediate key point data, and the initial key point data of the Nth frame can be used as the above-mentioned target face key data. Alternatively, the initial key point data of the (N - a)th frame, the initial key point data of the (N + a)th frame, and the initial key point data of the Nth frame can be weighted and averaged to calculate the above-mentioned target key point data. The specific calculation method can also be customized according to user requirements and is not specifically limited in this exemplary embodiment.
[0057] In an example embodiment of the present disclosure, in response to N being equal to M, the initial key point data of the (M - a)-th frame and the initial key point data of the M-th frame can be obtained, and the target key point data of the M-th frame can be calculated based on the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data. Specifically, the initial key point data of the (M - a)-th frame and the initial key point data of the M-th frame can be obtained by using the above-mentioned key point recognition sub-model, and then the third intermediate key point data of the M-th frame can be calculated based on the initial key point data of the (M - a)-th frame and the forward optical flow data from the (M - a)-th frame to the M-th frame. After that, the average value of the third intermediate key point data and the initial key point data of the M-th frame can be used as the target key point data of the M-th frame. The method for calculating the target key point data of the M-th frame can also be customized according to user requirements and is not specifically limited in this example embodiment.
[0058] In step S230, update the video recognition model according to the initial key point data and the target key point data;
[0059] In step S240, in response to the target key point data satisfying a preset condition, output the target key point data.
[0060] In an example embodiment of the present disclosure, updating the video recognition model according to the initial key point data and the target key point data may include steps S410 to S430, which are specifically as follows:
[0061] In step S410, calculate the error value between the initial key point data and the target key point data.
[0062] In this example embodiment, the error value between the initial key point data and the target key point data can be calculated. Specifically, the difference between the initial key point data and the target key point data can be taken and its absolute value can be used as the above-mentioned error value.
[0063] After calculating the above error value, it can be first determined whether the above target key point data satisfies the preset condition. If the above target key point data satisfies the preset condition, the above target key point data will be output. At this time, there is no need to update the above model. That is, when the above target key point data satisfies the preset condition, the update of the above video recognition model will be stopped. If the above target key point data does not satisfy the above preset condition, step S420 will be executed.
[0064] In this example embodiment, the above preset condition may be that the above error value is less than or equal to a second preset value. Among them, the value of the second preset value can be 2, 2.5, etc., or it can also be customized according to user requirements and is not specifically limited in this example embodiment.
[0065] In step S420, in response to the error value being less than or equal to the first preset value, the key point recognition sub-model is updated using the target key point data.
[0066] In the present exemplary embodiment, a first preset value can be set first. The value of the first preset value can be a positive integer such as 10, 9, etc., or a decimal such as 9.1, 9.5, etc., and can also be customized according to user requirements, and no specific limitation is made in the exemplary embodiment.
[0067] In the present exemplary embodiment, when the above-mentioned target key point data does not meet the above-mentioned preset conditions and the error value is less than or equal to the above-mentioned first preset value, it can be determined that the accuracy of the above-mentioned key point recognition sub-model is insufficient, and then the above-mentioned face key point recognition sub-model is updated using the above-mentioned target key point data.
[0068] In step S430, in response to the error value being greater than the first preset value, the optical flow calculation sub-model is updated using the initial key point data.
[0069] In the present exemplary embodiment, if the above-mentioned error value is greater than the above-mentioned first preset value, it is determined that the accuracy of the above-mentioned optical flow calculation sub-model is insufficient, and the above-mentioned optical flow calculation sub-model can be updated using the above-mentioned initial key point data.
[0070] In an exemplary embodiment of the present disclosure, the above steps S410 to S430 can be repeated until the above-mentioned target key point data meets the preset conditions, and finally the target key point data that meets the preset conditions is output. When the above-mentioned target key point data meets the preset conditions, the update of the model is stopped and the extraction of the video key point data is completed.
[0071] In an exemplary embodiment of the present disclosure, after the target key point data is obtained for the first time, it can be determined whether the target key point data meets the preset conditions, and the above-mentioned video recognition model is directly updated using the target key point data and the initial key point data. After the target key point data is obtained for the second time, it starts to be determined whether the above-mentioned target key point data meets the preset conditions. When the above-mentioned target key point data meets the preset conditions, the update of the above-mentioned video recognition model is stopped and the target key point data is output. If not, the above-mentioned video recognition model is continuously updated until the above-mentioned target key point data meets the preset conditions.
[0072] In an exemplary embodiment of the present disclosure, reference can be made to Figure 5As shown in the figure, a specific example implementation is used to illustrate the above video key point data extraction method. First, the initial key point data 503 of multiple frames of images 501 in the video can be obtained by using the key point recognition sub-model 502. Then, the above initial key point data 503 are all input into the above optical flow calculation sub-model 504 to obtain the optical flow vectors 505 between each frame of images. Then, the target key point data 506 are calculated according to the above initial key point data 503 and the optical flow vectors 505. After that, it is determined whether the above target key point data meet the preset conditions. If they meet, the target key point data are output. If they do not meet, the error value 507 between the above target key point data and the initial key point data is determined. If the above error value is less than or equal to the first preset value, the key point recognition sub-model 502 is updated by using the above target key point data. If it is greater than or equal to the first preset value, the optical flow calculation sub-model 504 is updated by using the above initial key point data.
[0073] Refer to Figure 6As shown above, taking the above a equal to 1 as an example, the above video face detection model will be described in detail. First, step S610 can be executed to decode the video to be recognized into frames, and then step S620 is executed to perform key point detection using the pre-trained key point recognition sub-model to obtain initial key point data; then step S630 is executed to calculate the optical flow vector for the input video frames. In the present exemplary embodiment, the optical flow vector may include a forward optical flow vector and a backward optical flow vector. The forward optical flow vector is the optical flow vector V(t-1,t) from the (t-1)th moment to the tth moment; the backward optical flow vector is the optical flow vector V(t+1,t) from the (t+1)th moment to the tth moment. Secondly, step S640 is executed to calculate the first intermediate key point data and the second intermediate key point data according to the optical flow vector for the key points, that is, the first intermediate key point data ldmk_det(t-1)+V(t-1,t) and the second intermediate key point data ldmk_det(t+1)+V(t+1,t). Then step S650 is executed to calculate the target key point data at the tth moment according to the first intermediate key point data and the second intermediate key point data ((ldmk_det(t-1)+V(t-1,t))+(ldmk_det(t+1)+V(t+1,t))+ldmk_det(t)) / 3; steps S610 to S650 are executed to filter the key points of all frames to obtain ldmk_f; then step S660 is executed to calculate the error value between the initial key point data and the target key point data; after obtaining the above error value, step S670 is executed to determine whether the error value meets a preset condition. If it meets, step S680 is executed to output the target key point data. If it does not meet, step S690 is executed. If the above error value is less than or equal to the first preset value, step S691 is executed to update the key point recognition sub-model using the above target key point data. If it is greater than or equal to the first preset value, step S692 is executed to update the above optical flow calculation sub-model using the above initial key point data. Specifically, by calculating the error value dif = ||ldmk_det - ldmk_f|| to determine whether the calculation result of the optical flow model is accurate. If the error value diff is less than the first preset value, it is considered that the optical flow calculation is accurate, and the key points ldmk_dif_small obtained by the optical flow calculation are used to replace the key points predicted by the model; otherwise, it is inaccurate, and the key points ldmk_dif_large predicted by the model are retained, and the above optical flow calculation sub-model is updated using the initial key point data.
[0074] In summary, in this exemplary embodiment, the video recognition model is updated according to the initial key point data and the target key point data, closely combining the optical flow data with the key points, without subsequent processing, reducing the complexity of video key point recognition. At the same time, in response to the target key point data meeting a preset condition, the target key point data is output, and the video recognition model is continuously updated during the recognition process until it meets the preset condition, which can improve the recognition accuracy of video key points. In an inaccurate scenario, the optical flow model improves the accuracy of optical flow model prediction by learning the form of coordinates output by the key point model, and improves the coordinate prediction stability of the key point detection model in an unsupervised mutual learning form between the optical flow model and the key point detection model.
[0075] It should be noted that the above drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above drawings do not indicate or limit the time sequence of these processes. Additionally, it is also easy to understand that these processes can be executed synchronously or asynchronously, for example, in multiple modules.
[0076] Further, as shown in Figure 7 In the embodiment of this example, a video key point recognition device 700 is further provided, including an acquisition module 710, a calculation module 720, an update module 730, and an output module 740. Among them:
[0077] The acquisition module 710 can be used to obtain the initial key point data of multiple frames of images in the video and the optical flow data between multiple frames of images by using the video recognition model.
[0078] In the embodiment of this example, the video recognition model includes: a key point recognition sub-model for obtaining the initial key point data in each frame of the video; an optical flow calculation sub-model for obtaining the optical flow data between multiple frames of the video.
[0079] The calculation module 720 can be used to determine the target key point data corresponding to each frame of the image by using the initial key point data and the optical flow data.
[0080] In an exemplary embodiment of the present disclosure, the video includes M frames of images. Calculating the target key point data corresponding to each frame of image by using the initial key point data and the optical flow data includes: in response to N being greater than 1 and less than M, obtaining the initial key point data of the (N - a)-th frame, the initial key point data of the N-th frame, and the initial key point data of the (N + a)-th frame; calculating the target key point data of the N-th frame according to at least one of the initial key point data of the (N - a)-th frame and the (N + a)-th frame, the initial key point data of the N-th frame, and the optical flow data; where a is a positive integer, and N - a is greater than or equal to 1, and N + a is less than or equal to M.
[0081] In this exemplary embodiment, the optical flow data includes forward optical flow data and backward optical flow data. Calculating the target key point data of the N-th frame according to at least one of the initial key point data of the (N - a)-th frame and the (N + a)-th frame, the initial key point data of the N-th frame, and the optical flow data includes: calculating the first intermediate key point data of the N-th frame according to the initial key point data of the (N - a)-th frame and the forward optical flow data from the (N - a)-th frame to the N-th frame; and / or calculating the second intermediate key point data of the N-th frame according to the initial key point data of the (N + a)-th frame and the backward optical flow data from the (N + a)-th frame to the N-th frame; calculating the target key point data of the N-th frame according to at least one of the second intermediate key point data of the N-th frame and the first intermediate key point data of the N-th frame, and the initial key point data of the N-th frame.
[0082] In an exemplary embodiment of the present disclosure, in response to N being equal to 1, the calculation module 720 uses the initial key point data of the first frame as the target key point data of the first frame; in response to N being equal to M, it obtains the initial key point data of the (M - a)-th frame and the initial key point data of the M-th frame, and calculates the target key point data of the M-th frame according to the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data. When calculating the target key point data of the M-th frame according to the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data, the target key point data of the M-th frame can be calculated according to the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data.
[0083] In an exemplary embodiment, the updating module 730 may be configured to update the video recognition model according to the initial key point data and the target key point data. Specifically, it calculates an error value between the initial key point data and the target key point data; in response to the error value being less than or equal to a first preset value, it updates the key point recognition sub-model by using the target key point data; in response to the error value being greater than the first preset value, it updates the optical flow calculation sub-model by using the initial key point data.
[0084] The output module 740 may be configured to output the target key point data in response to the target key point data satisfying a preset condition. Specifically, in response to the error value being less than or equal to a second preset value, it outputs the target key point data; wherein, the first preset value is greater than the second preset value.
[0085] The specific details of each module in the above device have been described in detail in the implementation manner of the method part. The details not disclosed can be referred to the implementation manner content of the method part, and thus will not be elaborated herein.
[0086] Next, taking Figure 8 the mobile terminal 800 in Figure 8 as an example, the structure of this electronic device will be described exemplarily. Those skilled in the art should understand that, except for the components specifically for mobile purposes,
[0087] such as Figure 8 shown, the mobile terminal 800 may specifically include: a processor 801, a memory 802, a bus 803, a mobile communication module 804, an antenna 1, a wireless communication module 805, an antenna 2, a display screen 806, a camera module 807, an audio module 808, a power module 809, and a sensor module 210.
[0088] The processor 801 may include one or more processing units. For example: the processor 801 may include an AP (Application Processor), a modulation and demodulation processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc. The video key point recognition method in this exemplary embodiment may be executed by the AP, GPU, or DSP. When the method involves neural network-related processing, it may be executed by the NPU.
[0089] An encoder can encode (i.e., compress) an image or video. For example, it can encode a target image into a specific format to reduce the data size for easier storage or transmission. A decoder can decode (i.e., decompress) the encoded data of an image or video to restore the image or video data. For example, it can read the encoded data of a target image, decode it through the decoder to restore the data of the target image, and then perform related processing for video key point recognition on this data. The mobile terminal 800 can support one or more encoders and decoders. In this way, the mobile terminal 800 can process images or videos in multiple encoding formats, such as: image formats like JPEG (Joint Photographic Experts Group), PNG (Portable Network Graphics), BMP (Bitmap), etc., and video formats like MPEG (Moving Picture Experts Group) 1, MPEG2, H.263, H.264, HEVC (High Efficiency Video Coding), etc.
[0090] The processor 801 can form a connection with the memory 802 or other components through the bus 803.
[0091] The memory 802 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 801 executes various functional applications and data processing of the mobile terminal 800 by running the instructions stored in the memory 802. The memory 802 can also store application data, such as storing files like images and videos.
[0092] The communication function of the mobile terminal 800 can be implemented through a mobile communication module 804, antenna 1, a wireless communication module 805, antenna 2, a modulation and demodulation processor, and a baseband processor, etc. Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 804 can provide 2G, 3G, 4G, 5G, etc. mobile communication solutions applied to the mobile terminal 800. The wireless communication module 805 can provide wireless communication solutions such as wireless local area network, Bluetooth, and near-field communication applied to the mobile terminal 800.
[0093] The display screen 806 is used to implement display functions, such as displaying user interfaces, images, videos, etc. The camera module 807 is used to implement shooting functions, such as shooting images, videos, etc. The audio module 808 is used to implement audio functions, such as playing audio, collecting voices, etc. The power module 809 is used to implement power management functions, such as charging the battery, powering the device, monitoring the battery status, etc. The sensor module 810 may include a depth sensor 8101, a pressure sensor 8102, a gyroscope sensor 8103, a barometric pressure sensor 8104, etc., to implement corresponding sensing and detection functions.
[0094] Those skilled in the art of the present disclosure can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuits", "modules", or "systems" here.
[0095] The exemplary embodiments of the present disclosure also provide a computer-readable storage medium, on which a program product capable of implementing the methods described above in this specification is stored. In some possible embodiments, various aspects of the present disclosure can also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Methods" section above in this specification.
[0096] It should be noted that the computer-readable medium shown in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0097] In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0098] In addition, the program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0099] Those skilled in the art will readily think of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0100] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. A method for video key point recognition, characterized in that, it includes: obtaining initial key point data of multiple frames of images in the video and optical flow data between multiple frames of images by using a video recognition model; determining target key point data corresponding to each frame of image by using the initial key point data and the optical flow data; updating the video recognition model according to the initial key point data and the target key point data; responding to the target key point data meeting a preset condition and outputting the target key point data; the video includes M frames of images, and calculating the target key point data corresponding to each frame of image by using the initial key point data and the optical flow data includes: when N is greater than 1 and less than M, obtaining the initial key point data of the (N - a)-th frame, the initial key point data of the N-th frame, and the initial key point data of the (N + a)-th frame; calculating the target key point data of the N-th frame according to at least one of the initial key point data of the (N - a)-th frame and the (N + a)-th frame, the initial key point data of the N-th frame, and the optical flow data; wherein, a is a positive integer, and N - a is greater than or equal to 1, and N + a is less than or equal to M.
2. The method according to claim 1, characterized in that, the optical flow data includes forward optical flow data and backward optical flow data, and calculating the target key point data of the N-th frame according to at least one of the initial key point data of the (N - a)-th frame and the (N + a)-th frame, the initial key point data of the N-th frame, and the optical flow data includes: calculating first intermediate key point data of the N-th frame according to the initial key point data of the (N - a)-th frame and the forward optical flow data from the (N - a)-th frame to the N-th frame; and / or calculating second intermediate key point data of the N-th frame according to the initial key point data of the (N + a)-th frame and the backward optical flow data from the (N + a)-th frame to the N-th frame; calculating the target key point data of the N-th frame according to at least one of the second intermediate key point data of the N-th frame and the first intermediate key point data of the N-th frame, and the initial key point data of the N-th frame.
3. The method according to claim 1, characterized in that, calculating the target key point data corresponding to each frame of image by using the initial key point data and the optical flow data includes: when N is equal to 1, taking the initial key point data of the first frame as the target key point data of the first frame; when N is equal to M, obtaining the initial key point data of the (M - a)-th frame and the initial key point data of the M-th frame, and calculating the target key point data of the M-th frame according to the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data.
4. The method according to claim 3, characterized in that, the optical flow data includes forward optical flow data and backward optical flow data, and calculating the target key point data of the M-th frame according to the initial key point data of the (M - a)-th frame, the initial key point data of the M-th frame, and the optical flow data includes: Calculate the third intermediate key point data of the Mth frame based on the initial key point data of the (M - a)th frame and the forward optical flow data from the (M - a)th frame to the Mth frame; Calculate the target key point data of the Mth frame based on the third intermediate key point data and the initial key point data of the Mth frame.
5. The method according to claim 1, wherein, the video recognition model includes: a key point recognition sub-model for obtaining the initial key point data in each frame image of the video; an optical flow calculation sub-model for obtaining the optical flow data between multiple frame images in the video.
6. The method according to claim 5, wherein, updating the key point recognition sub-model and the optical flow calculation sub-model according to the initial key point data and the target key point data includes: calculating an error value between the initial key point data and the target key point data; in response to the error value being less than or equal to a first preset value, updating the key point recognition sub-model using the target key point data; in response to the error value being greater than the first preset value, updating the optical flow calculation sub-model using the initial key point data.
7. The method according to claim 6, wherein, outputting the target key point data in response to the target key point data satisfying a preset condition includes: in response to the error value being less than or equal to a second preset value, outputting the target key point data; wherein, the first preset value is greater than the second preset value.
8. A video key point recognition device, wherein, comprises: an acquisition module for obtaining the initial key point data of multiple frame images in a video and the optical flow data between multiple frame images by using a video recognition model; a calculation module for determining the target key point data corresponding to each frame image by using the initial key point data and the optical flow data; an update module for updating the video recognition model according to the initial key point data and the target key point data; an output module for outputting the target key point data in response to the target key point data satisfying a preset condition; the video includes M frame images, and calculating the target key point data corresponding to each frame image by using the initial key point data and the optical flow data includes: in response to N being greater than 1 and less than M, obtaining the initial key point data of the (N - a)th frame, the initial key point data of the Nth frame, and the initial key point data of the (N + a)th frame; calculating the target key point data of the Nth frame according to at least one of the initial key point data of the (N - a)th frame and the (N + a)th frame, the initial key point data of the Nth frame, and the optical flow data; where a is a positive integer, and N - a is greater than or equal to 1, and N + a is less than or equal to M.
9. A computer-readable storage medium, on which a computer program is stored, wherein, the computer program, when executed by a processor, implements the video key point recognition method according to any one of claims 1 to 7.
10. An electronic device, wherein, comprises: one or more processors; and A memory for storing one or more programs, which when executed by the one or more processors, cause the one or more processors to implement the video key point recognition method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for processing video, equipment and storage medium
CN110148158A
Human body posture acquisition method and device, electronic equipment and storage medium
CN113569781A