Video recognition detection model generation method and apparatus, and computer device
By introducing a jitter loss function into the video recognition and detection network to adjust the model parameters, the problem of reduced accuracy and stability in video recognition and detection is solved, achieving high-precision and stable video recognition and detection.
Patent Information
- Application Number
- CN202110601898.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-31
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2041-05-31
AI Technical Summary
When deep learning-based video recognition and detection networks are applied to videos or continuous images, their accuracy and stability are greatly reduced due to the influence of image jitter during prediction.
By inputting the frame image sequence into a preset first network model for processing, a predicted heatmap is output, and the parameters of the network model are adjusted according to the jitter loss function until the training conditions are met, thus generating a trained video recognition and detection model.
It improves the accuracy and stability of video recognition and detection, and eliminates the jitter effect of the predicted frame image through a simple training process.
Smart Images

Figure CN115424158B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image processing, and particularly relates to a video recognition detection model generation method and device and computer equipment. BACKGROUND
[0002] Facial landmark detection (also known as face alignment) is an active branch of computer vision research. Facial landmark detection is a key step in the field of face recognition and analysis, and it is a prerequisite and breakthrough for other face-related problems such as automatic face recognition, expression analysis, three-dimensional face reconstruction, and three-dimensional animation. In recent years, deep learning methods have been successfully applied to many fields such as image recognition and analysis, speech recognition, and natural language processing due to their automatic learning and continuous learning capabilities, and have brought significant improvements in these areas. Video recognition detection networks based on deep learning have achieved excellent performance on static images. However, when video recognition detection networks based on deep learning are applied to videos or continuous images, their video recognition detection accuracy and stability are greatly reduced due to the influence of predicted image jitter.
[0003] Therefore, the prior art still needs to be improved and developed. SUMMARY
[0004] In view of the above defects of the prior art, the embodiments of the present application provide a video recognition detection model generation method and device and computer equipment, aiming to solve the problem that the video recognition detection accuracy and stability of video recognition detection networks based on deep learning are greatly reduced when they are applied to videos or continuous images due to the influence of predicted image jitter.
[0005] The technical solutions adopted by the present application to solve the problems are as follows:
[0006] In a first aspect, the embodiments of the present application provide a video recognition detection model generation method, comprising:
[0007] inputting a frame image sequence into a preset first network model for processing to output a predicted heat map corresponding to the frame image sequence; wherein the frame image sequence is obtained by preprocessing frame images in training data, and the training data further includes a real label, the real label representing a real heat map corresponding to the frame images;
[0008] determining a jitter loss function according to the predicted heat map and the real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images;
[0009] adjusting the parameters of the preset first network model according to the jitter loss function, and continuing to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, to obtain a trained video recognition detection model.
[0010] In a second aspect, the embodiments of the present application further provide a video recognition detection method, comprising:
[0011] obtaining a frame image sequence to be processed;
[0012] inputting the frame image sequence to be processed into a trained video recognition detection model for processing, and outputting a target heat map.
[0013] In a third aspect, the embodiments of the present application provide a video recognition detection model generation device, comprising:
[0014] a predicted heat map acquisition module configured to input a frame image sequence into a preset first network model for processing, and output a predicted heat map corresponding to the frame image sequence; wherein the frame image sequence is obtained by preprocessing frame images in training data, and the training data further comprises a real label, the real label representing a real heat map corresponding to the frame images;
[0015] a jitter loss function determination module configured to determine a jitter loss function according to the predicted heat map and the real label; wherein the jitter loss function is used to eliminate jitter effects between adjacent frame images;
[0016] a video recognition detection model acquisition module configured to adjust parameters of the preset first network model according to the jitter loss function, and continue to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, and obtain a trained video recognition detection model.
[0017] In a fourth aspect, the embodiments of the present application further provide a video recognition detection device, comprising:
[0018] a frame image sequence acquisition module configured to obtain a frame image sequence to be processed;
[0019] a target heat map acquisition module configured to input the frame image sequence to be processed into a trained video recognition detection model for processing, and output a target heat map.
[0020] In a fifth aspect, the embodiments of the present application provide a computer device, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the video recognition detection model generation method of any one of the above aspects when executing the computer program.
[0021] In a sixth aspect, the embodiments of the present application further provide a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the video recognition detection model generation method of any one of the above aspects.
[0022] The beneficial effects of the present application are: firstly, the frame image sequence is input to the preset first network model for processing, and a predicted heat map corresponding to the frame image sequence is output, and the heat map accuracy can be greatly improved by training the first network model with little effort, wherein the frame image sequence is obtained by preprocessing the frame images in the training data, and the training data further includes a real label, and the real label represents a real heat map corresponding to the frame image; then, a jitter loss function is determined according to the predicted heat map and the real label, and the obtained jitter loss function prepares for subsequent elimination of the jitter influence of the predicted frame image; finally, the parameters of the preset first network model are adjusted according to the jitter loss function, and the step of inputting the frame image sequence to the preset first network model for processing is continued to be performed until the preset training condition is met, and a trained video recognition detection model is obtained. The video recognition detection model obtained by the method can improve the recognition detection accuracy through simple training, and the jitter influence of the predicted frame image can be eliminated through the jitter loss function, so that the accuracy and stability of the video recognition detection task are improved. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0024] Figure 1 The video recognition detection model generation method flowchart provided by the embodiment of the present application.
[0025] Figure 2 The first network structure diagram provided by the embodiment of the present application.
[0026] Figure 3 The flowchart of the video recognition detection method provided by the embodiment of the present application.
[0027] Figure 4 The functional module diagram of the video recognition detection model generation device provided by the embodiment of the present application.
[0028] Figure 5 The functional module diagram of the video recognition detection device provided by the embodiment of the present application.
[0029] Figure 6 The internal structure principle block diagram of the computer equipment provided by the embodiment of the present application. DETAILED DESCRIPTION
[0030] The application discloses a video recognition detection model generation method and a computer device. In order to make the purpose, technical scheme and effect of the application more clear and definite, the application is further described in detail below with reference to the drawings and examples. It should be understood that the specific examples described herein are only used to explain the application and not to limit the application.
[0031] Those skilled in the art can understand that the singular forms "a," "an," and "the" used herein include plural forms unless specifically stated otherwise. It should be further understood that the use of the term "includes" in the specification of the application means that a feature, integer, step, operation, element, and / or component exists, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there can be intermediate elements. In addition, "connected" or "coupled" used herein can include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any single unit and all combinations of the associated listed items.
[0032] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as that generally understood by those skilled in the art to which the application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined as such.
[0033] In the prior art, the video recognition detection network based on deep learning greatly reduces the video recognition detection accuracy and stability when applied to video or continuous images due to the influence of predicted image jitter.
[0034] Examples
[0035] The static picture recognition technology is very mature, but the precision and stability of video recognition detection are not high, and the video-based detector is easily affected by time tracker drift, if the time tracker tracks the current frame incorrectly, the subsequent video frame recognition detection will be affected, to solve the two problems, the influence of the predicted frame image jitter needs to be eliminated, therefore, the embodiment of the application provides a video recognition detection model generation method: first, input the frame image sequence into a preset first network model for processing, output the predicted heat map corresponding to the frame image sequence; only a little effort is needed to train the first network model to greatly improve the heat map precision, wherein the frame image sequence is obtained by preprocessing the frame image in the training data, the training data also includes a real label, and the real label represents the real heat map corresponding to the frame image; then, according to the predicted heat map and the real label, a jitter loss function is determined, and the jitter loss function obtained is prepared for subsequent elimination of the jitter influence of the predicted frame image; finally, the parameters of the preset first network model are adjusted according to the jitter loss function, and the step of inputting the frame image sequence into the preset first network model for processing is continued to be performed until the preset training condition is met, and a trained video recognition detection model is obtained, the video recognition detection model improves the recognition detection precision through simple training, and the jitter influence of the predicted frame image can be eliminated through the jitter loss function, so that the precision and stability of the video recognition detection task are improved.
[0036] Exemplary method
[0037] The embodiment provides a video recognition detection model generation method, which can be applied to a computer device for video image processing. Figure 1 As shown in the figure, the method comprises the following steps:
[0038] In step S100, a frame image sequence is input into a preset first network model for processing, and a predicted heat map corresponding to the frame image sequence is output; wherein the frame image sequence is obtained by preprocessing the frame image in the training data, and the training data also includes a real label, and the real label represents the real heat map corresponding to the frame image;
[0039] Specifically, to implement the video recognition detection task, the video needs to be processed, and for the processed video, the video is a series of frame images, therefore, in the training process, the preset first network model needs to be trained by inputting the video stream, that is, the sequence of frame images. In practice, the original frame image can be processed to obtain the sequence of frame images, and the processing method can be moving the pixels of the original frame image, or processing the frame image based on the deep learning method. In this embodiment, the frame image in the training data is preprocessed to obtain. For a model, it will be trained, and the training will involve training data. In this embodiment, the training data includes frame images and real labels. The real label refers to the expected target expected to be obtained after inputting a known input data into a network model for processing. In this embodiment, the real label represents the real heat map corresponding to the frame image. The sequence of frame images is input into the first model, and the model will have an output, that is, the predicted heat map corresponding to the sequence of frame images.
[0040] In one implementation manner of the embodiment of the present application, before the sequence of frame images is input into the preset first network model for processing and the predicted heat map corresponding to the sequence of frame images is output, the method further includes the following steps:
[0041] According to the frame image in the training data, the frame image pixel and the key point pixel in the frame image in the training data are obtained.
[0042] The frame image pixel and the key point pixel are moved in the horizontal direction or the vertical direction according to the preset offset to obtain the sequence of frame images corresponding to the frame image in the training data.
[0043] Specifically, before the sequence of frame images is input into the preset first network model for processing and the predicted heat map corresponding to the sequence of frame images is output, when each frame image in the video is recognized and detected, the frame image pixel, that is, the pixel information of each frame image, can be obtained. The key point in the frame image is recognized through each frame image, and the key point pixel, that is, the pixel information of the key point, can be obtained. Finally, the frame image pixel and the key point pixel can be moved in the horizontal direction or the vertical direction according to the preset offset to obtain the sequence of frame images corresponding to the frame image. In practice, the offset of the movement of the frame image pixel in the horizontal direction or the vertical direction can be the same as the offset of the movement of the key point pixel in the horizontal direction or the vertical direction, or can be different. The sequence of frame images corresponding to the frame image can be obtained by moving the frame image pixel and the key point pixel in the horizontal direction, or can be obtained by moving the frame image pixel and the key point pixel in the vertical direction, or can be obtained by moving the frame image pixel and the key point pixel in the horizontal direction and the vertical direction at the same time.
[0044] In an implementation form of the embodiment of the present application, the first network model comprises a pre-trained backbone network and a long short-term memory network; the frame image sequence is input to the preset first network model for processing, and a predicted heat map corresponding to the frame image sequence is output, comprising:
[0045] In step S101, the frame image sequence is input to the pre-trained backbone network for processing, and an initial predicted heat map is output.
[0046] Specifically, as shown in the figure, Figure 2 The pre-trained backbone network can be replaced by any compatible network supporting heat map regression, such as HRNet and HourGlass, etc. The preset pre-trained backbone network is a trained network, which serves as a feature extractor to extract the feature map of the frame image. The frame image sequence I t is input to the preset pre-trained backbone network, and an unstable predicted heat map o t is obtained. Since the preset pre-trained backbone network is a trained network, it does not need to be trained a lot, but only needs to be fine-tuned to obtain an initial heat map o t . The initial heat map o t is unstable and is prepared for further extraction of heat map by the subsequent network, which can greatly improve the recognition and detection accuracy of the frame image.
[0047] In step S102, the initial predicted heat map is input to the long short-term memory network for processing, and a predicted heat map corresponding to the frame image sequence is output.
[0048] Specifically, after obtaining the initial predicted heat map o t , the initial predicted heat map o t is input to the preset convolutional long short-term memory network to obtain the predicted heat map. The convolutional long short-term memory network is a unit in which an S-shaped activation is applied to the linear combination of its input, which is replaced by a storage unit. Each storage unit is associated with an input gate, an output gate and an internal state that is sent to itself without interference across time steps, so that the convolutional long short-term memory network can not only obtain the spatial semantic information of the frame image, but also obtain the temporal semantic information of the frame image. When the preset convolutional long short-term memory network is trained, the initial predicted heat map o t is input to the convolutional long short-term memory network, and the obtained predicted heat map u t is a stable heat map.
[0049] After obtaining the predicted heat map u t , the following steps can be performed as shown in the figure, Figure 1The method comprises the following steps: step S200, determining a jitter loss function according to a predicted heat map and a real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images;
[0050] Specifically, since the existing loss function for heat map regression needs to calculate the pixel-level error between the predicted value and the real value, this index is very good on static images. However, as a loss function for measuring the pixel error of a single frame image, it cannot sufficiently suppress the negative effects of jittering of the key points of the predicted frame image around the target points, resulting in obvious jittering of the prediction of the video recognition detection model trained by this loss function. Therefore, the embodiment of the present application proposes a jitter loss function, which is used to eliminate the jitter effect between adjacent frame images, and the jitter loss function can be calculated according to the difference relationship between the predicted heat map and the real label. Correspondingly, the determination of the jitter loss function according to the predicted heat map and the real label comprises the following steps:
[0051] S201, determining a first difference value according to the predicted heat map and the real label of the current frame;
[0052] Specifically, the predicted heat map u of the current frame is subtracted from the real label of the current frame t to obtain the predicted frame image difference of the current frame between the predicted heat map of the current frame and the real label of the current frame, that is, the first difference value, for example:
[0053] S202, determining a modulation function according to the first difference value, the predicted heat map and the real label;
[0054] In the embodiment, step S202 comprises the following steps:
[0055] determining a second difference value according to the predicted heat map and the real label of the previous frame;
[0056] determining an adjacent frame predicted value error according to the first difference value and the second difference value;
[0057] determining an adjacent frame real value deviation according to the real label of the current frame and the real label of the previous frame;
[0058] determining the modulation function according to the adjacent frame predicted value error and the adjacent frame real value deviation.
[0059] Specifically, the predicted heat map u of the previous frame is subtracted from the real label of the previous frame t-1 to obtain the second difference value, for example: Then, the first difference value is subtracted from the second difference value to obtain the adjacent frame predicted value error, for example: e t -e t-1 Thus, when the discrepancy between the predicted values of two adjacent frames exceeds a certain threshold, jitter occurs. The true label of the current frame is then... Subtract the real label from the previous frame The deviation between the true values of adjacent frames is obtained, for example: Finally, based on the prediction error between adjacent frames and the deviation between the actual values in adjacent frames, the modulation function is obtained, for example, the modulation function. The upper limit of the modulation function is the weight threshold w, which is a large and fixed weight threshold w. The purpose of w is to enhance the jitter loss so that the deviation of the true value converges as soon as possible. The weight threshold w can be set to 5 to avoid gradient explosion. The denominator of the modulation function includes the regularization parameter ξ to avoid the occurrence of singular values during training.
[0060] S203. Determine the pixel loss function based on the first difference;
[0061] In this embodiment, step S203 specifically includes the following steps: calculating the square of the first difference to obtain the squared difference value; adding a preset jitter threshold to the squared difference value to obtain the sum of squares value; dividing the squared difference value by the sum of squares value to obtain the pixel loss function.
[0062] Specifically, pixel loss function The choice of loss function is crucial. Since the network of this invention pre-assumes a pre-trained backbone network, this pre-trained backbone network requires fine-tuning rather than training from scratch. The first network focuses on training a pre-defined convolutional long short-term memory network to address the jitter detection problem in video frame images, while simultaneously determining the weights of the pre-trained backbone network. This invention does not choose L2 and smoothed L1 loss functions because their gradients near the origin are small. While Wing and L1 loss functions alleviate the small gradient problem, they introduce discontinuities at the origin, increasing the difficulty of pre-training the backbone network. The Awing loss function is an improvement on the Wing loss function, capable of handling smaller errors while ensuring a continuous slope; however, the Awing loss function is designed to handle outliers, not jitter. Therefore, this invention proposes a jitter loss function, which requires first obtaining the pixel loss function. In practice, the first difference is squared to obtain the squared difference value; then the squared difference is added to a preset jitter threshold Θ to obtain the sum of squares value; the squared difference value is divided by the sum of squares value to obtain the pixel loss function; for example, the pixel loss function is: ,when Increasing the gradient to Θ / 2 causes the pixel loss function gradient to reach an extreme value, and then gradually decreases to zero. This makes the optimizer quite robust in handling outliers, thus effectively handling small errors caused by jittery key points in video frames.
[0063] S204, determining a jitter loss function according to the pixel loss function and the modulation function.
[0064] In the embodiment, step S204 includes the following step: calculating the product of the pixel loss function and the modulation function to obtain the jitter loss function.
[0065] Specifically, the pixel loss function is multiplied by the modulation function to obtain the jitter loss function, for example: where u t and are the predicted value and the true value of the t-th frame respectively. The jitter loss function not only considers the prediction error of each frame, but also considers the jitter effect of adjacent frames. The jitter loss is linearly dependent on the c t-1,t standardized e t -e t-1 , and is also constrained by the weight threshold w.
[0066] After obtaining the jitter loss function, the following steps can be performed as in Figure 1 Step S300, adjusting the parameters of the preset first network model according to the jitter loss function, and continuing to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met to obtain a trained video recognition detection model.
[0067] In the embodiment, after obtaining the jitter loss function, the parameters of the first network model can be adjusted according to the jitter loss function. For example, when the jitter loss function is large, it means that the network model is far from the trained target, and at this time, the parameters such as weights of the first network model need to be adjusted to a larger value to quickly converge the jitter loss function; when the jitter loss function is small, it means that the network model is very close to the trained target, and at this time, only the parameters such as weights in the first network model need to be fine-tuned to converge the jitter loss function. While adjusting the parameters of the first network model according to the jitter loss function, the frame image sequence is continuously input into the first network model for training other frame image sequences, so that the first network model obtains complete semantic information of the frame image sequence, thereby obtaining a trained video recognition detection model. When the training of the first network model is completed, the pre-training backbone network and the convolutional long short-term memory network are also trained.
[0068] In an implementation manner of the embodiment, the initial heat map o tand the predicted heat map are probability maps, which describe the probability of the key point in the frame image appearing in the corresponding position in the image, and the predicted key point is usually located at the center of the heat map. In the prior art, the key point is obtained by selecting the maximum point (for example, the argmax method and the interpolation method), but the argmax method cannot achieve sub-pixel accuracy, and the interpolation method can well process static images, but when processing video frame images, it is covered by random noise and motion blur, because the interpolation method is very sensitive to small errors of background pixels, which increases the difficulty of estimating the center of the heat map. Therefore, the embodiment of the present application proposes a probability density centering (PDC) algorithm to calculate the center of the initial heat map o t and the predicted heat map u t The PDC algorithm uses global information of the heat map to document the prediction result, and in addition, the present application filters out some pixel values of the key point below a preset threshold to eliminate small error interference caused by jitter of background pixels, and obtains the center of the heat map, i.e. the result centroid, by integrating the probability density of each pixel in the heat map.
[0069] In another implementation manner of the embodiment, in the process of fine-tuning the pre-trained backbone network, an Adam optimizer is used, and the initial learning rate is 1e -4 or 1e -5 .
[0070] As shown in Figure 3 , the embodiment of the present application provides a video recognition detection method, which comprises:
[0071] A100, acquiring a frame image sequence to be processed;
[0072] Specifically, the frame image sequence can be collected from a network or a data center, wherein the frame image sequence can also be a video stream, and the frame image sequence comprises a plurality of frame images.
[0073] A200, inputting the frame image sequence to be processed into a trained video recognition detection model for processing, and outputting a target heat map.
[0074] Specifically, the video recognition detection model comprises a pre-trained backbone network and a convolutional long short-term memory network, the frame image sequence to be processed is first input into the trained pre-trained backbone network to obtain an unstable initial predicted heat map, at this time, because the frame image is extracted once, the quality of subsequent heat map processing can be improved, and then the initial predicted heat map is input into the trained convolutional long short-term memory network, and because the convolutional long short-term memory network in the video recognition detection model can eliminate the jitter influence of the frame image in the video stream, a stable heat map can be obtained.
[0075] As shown in Figure 4As shown in the figure, this embodiment of the invention provides a video recognition and detection model generation device, including a prediction heatmap acquisition module 401, a jitter loss function determination module 402, and a video recognition and detection model acquisition module 403, wherein:
[0076] The predicted heatmap acquisition module 401 is used to input the frame image sequence into a preset first network model for processing and output the predicted heatmap corresponding to the frame image sequence; wherein, the frame image sequence is obtained by preprocessing the frame images in the training data, and the training data also includes real labels, which represent the real heatmap corresponding to the frame images.
[0077] The jitter loss function determination module 402 is used to determine the jitter loss function based on the predicted heatmap and the true label; wherein, the jitter loss function is used to eliminate the jitter effect between adjacent frame images;
[0078] The video recognition and detection model acquisition module 403 is used to adjust the parameters of the preset first network model according to the jitter loss function, and continue to execute the step of inputting the frame image sequence into the preset first network model for processing until the preset training conditions are met, so as to obtain the trained video recognition and detection model.
[0079] In this embodiment, the frame image sequence is input into a preset first network model by the predictive heatmap acquisition module 401 for processing, and the predicted heatmap corresponding to the frame image sequence is output. The accuracy of the heatmap can be greatly improved by only spending a small amount of effort to train the first network model. The frame image sequence is obtained by preprocessing the frame images in the training data, which also includes real labels, which represent the real heatmaps corresponding to the frame images. Then, the jitter loss function determination module 402 determines the jitter loss function based on the predicted heatmap and the real labels. The obtained jitter loss function prepares for the subsequent elimination of the jitter effect of the predicted frame images. The jitter loss function is used to eliminate the jitter effect between adjacent frame images. Finally, the video recognition and detection model acquisition module 403 adjusts the parameters of the preset first network model according to the jitter loss function and continues to execute the step of inputting the frame image sequence into the preset first network model for processing until the preset training conditions are met, and a trained video recognition and detection model is obtained. The video recognition and detection model obtained by this method can improve the accuracy of recognition and detection through simple training. The jitter loss function can eliminate the jitter effect of the predicted frame image, thereby improving the accuracy and stability of the video recognition and detection task.
[0080] like Figure 5 As shown in the figure, this embodiment of the invention provides a video recognition and detection device, which includes a frame image sequence acquisition module 501 and a target heatmap acquisition module 502, wherein:
[0081] The frame image sequence acquisition module 501 is used to acquire the frame image sequence to be processed.
[0082] The target heatmap acquisition module 502 is used to input the frame image sequence to be processed into the trained video recognition and detection model for processing and output the target heatmap.
[0083] In this embodiment, the frame image sequence to be processed is obtained by the frame image sequence acquisition module 501. Since features are extracted from the frame images first, the quality of subsequent heatmap processing can be improved. The frame image sequence to be processed is input into the trained video recognition and detection model for processing by the target heatmap acquisition module 502, and the target heatmap is output. Since the convolutional long short-term memory network in the video recognition and detection model can eliminate the jitter of the frame images in the video stream, a stable heatmap can be obtained by processing the frame image sequence to be processed by the trained video recognition and detection model.
[0084] Based on the above embodiments, the present invention also provides a computer device, the principle block diagram of which can be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, display screen, and temperature sensor connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video recognition and detection model generation method. The display screen can be an LCD screen or an e-ink screen. The temperature sensor is pre-installed inside the computer device to detect the operating temperature of the internal components.
[0085] Those skilled in the art will understand that Figure 6 The schematic diagrams shown are merely block diagrams of some structures related to the present invention and do not constitute a limitation on the computer device to which the present invention is applied. Specific computer devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0086] In one embodiment, a computer device is provided, the computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing instructions to perform the following operations when executing the computer program:
[0087] inputting the frame image sequence into a preset first network model for processing, and outputting a predicted heat map corresponding to the frame image sequence; wherein the frame image sequence is obtained by preprocessing frame images in training data, and the training data further comprises a real label, the real label representing a real heat map corresponding to the frame images;
[0088] determining a jitter loss function according to the predicted heat map and the real label; wherein the jitter loss function is used to eliminate jitter effects between adjacent frame images;
[0089] adjusting parameters of the preset first network model according to the jitter loss function, and continuing to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, to obtain a trained video recognition detection model.
[0090] or
[0091] obtaining a frame image sequence to be processed;
[0092] inputting the frame image sequence to be processed into the trained video recognition detection model for processing, and outputting a target heat map.
[0093] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM).
[0094] To sum up, the application discloses a video recognition detection model generation method and device and computer equipment, and the method comprises the following steps: firstly, inputting a frame image sequence into a preset first network model for processing, and outputting a predicted heat map corresponding to the frame image sequence, so that the heat map precision can be greatly improved by training the first network model with little effort, wherein the frame image sequence is obtained by preprocessing frame images in training data, and the training data further comprises a real label, and the real label represents a real heat map corresponding to the frame images; then, determining a jitter loss function according to the predicted heat map and the real label, so as to prepare for subsequent elimination of the jitter influence of the predicted frame image; finally, adjusting the parameters of the preset first network model according to the jitter loss function, and continuing to execute the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, so as to obtain a trained video recognition detection model. The video recognition detection model obtained by the method can improve the recognition detection precision through simple training, and the jitter influence of the predicted frame image can be eliminated through the jitter loss function, so that the precision and stability of the video recognition detection task are improved.
[0095] Based on the above embodiments, the application discloses a video recognition detection model generation method, and it should be understood that the application is not limited to the above examples, and for those skilled in the art, the above description can be improved or changed, and all these improvements and changes should belong to the protection scope of the claims of the application.
Claims
1. A method for generating a video recognition detection model, characterized in that, The method comprises the following steps: inputting a frame image sequence into a preset first network model for processing, and outputting a predicted heat map corresponding to the frame image sequence; wherein the frame image sequence is obtained by preprocessing frame images in training data, and the training data further comprises a real label, wherein the real label represents a real heat map corresponding to the frame image; determining a jitter loss function according to the predicted heat map and the real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images; adjusting the parameters of the preset first network model according to the jitter loss function, and continuing to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, thereby obtaining a trained video recognition detection model; determining a jitter loss function according to the predicted heat map and the real label, comprising: determining a first difference value according to the predicted heat map and the real label of the current frame; determining a modulation function according to the first difference value, the predicted heat map and the real label; the modulation function is provided with an upper limit, and the upper limit is a preset weight threshold; the weight threshold is used to enhance the jitter loss; determining a pixel loss function according to the first difference value; determining a jitter loss function according to the pixel loss function and the modulation function.
2. The method of claim 1, wherein, Before the step of inputting the frame image sequence into the preset first network model for processing, and outputting a predicted heat map corresponding to the frame image sequence, the method further comprises: obtaining frame image pixels and key point pixels in the frame images in the training data according to the frame images in the training data; moving the frame image pixels and the key point pixels according to a preset offset in the horizontal direction or the vertical direction to obtain a frame image sequence corresponding to the frame images in the training data.
3. The method of claim 2, wherein, The first network model comprises a pre-trained backbone network and a long short-term memory network; the step of inputting the frame image sequence into the preset first network model for processing, and outputting a predicted heat map corresponding to the frame image sequence, comprises: inputting the frame image sequence into the pre-trained backbone network for processing, and outputting an initial predicted heat map; inputting the initial predicted heat map into the long short-term memory network for processing, and outputting a predicted heat map corresponding to the frame image sequence.
4. The method of claim 1, wherein, The step of determining a modulation function according to the first difference value, the predicted heat map and the real label, comprises: determining a second difference value according to the predicted heat map and the real label of the previous frame; determining an adjacent frame predicted value error according to the first difference value and the second difference value; determining an adjacent frame real value deviation according to the real label of the current frame and the real label of the previous frame; determining a modulation function according to the adjacent frame predicted value error and the adjacent frame real value deviation.
5. The method of claim 4, wherein, The step of determining a pixel loss function according to the first difference value, comprises: calculating the square of the first difference value to obtain a difference value square value; adding a preset jitter threshold to the difference value square value to obtain a sum square value; dividing the difference value square value by the sum square value to obtain a pixel loss function.
6. The method of claim 5, wherein, The determining of the jitter loss function according to the pixel loss function and the modulation function comprises: calculating the product of the pixel loss function and the modulation function to obtain the jitter loss function.
7. A video recognition detection method, characterized by, The method comprises: obtaining a frame image sequence to be processed; inputting the frame image sequence to be processed into a trained video recognition detection model for processing, and outputting a target heat map; The generation of the video recognition detection model comprises: determining a jitter loss function according to a predicted heat map and a real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images; The determining of the jitter loss function according to the predicted heat map and the real label comprises: determining a first difference value according to the predicted heat map and the real label of a current frame; determining a modulation function according to the first difference value, the predicted heat map and the real label; the modulation function is provided with an upper limit, and the upper limit is a preset weight threshold; the weight threshold is used to enhance the jitter loss; determining a pixel loss function according to the first difference value; determining the jitter loss function according to the pixel loss function and the modulation function. 8.A video recognition detection model generation apparatus, characterized by comprising: The method comprises: a predicted heat map acquisition module, configured to input a frame image sequence into a preset first network model for processing, and output a predicted heat map corresponding to the frame image sequence; wherein the frame image sequence is obtained by preprocessing frame images in training data; the training data further comprises a real label, and the real label represents a real heat map corresponding to the frame image; a jitter loss function determination module, configured to determine a jitter loss function according to the predicted heat map and the real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images; a video recognition detection model acquisition module, configured to adjust parameters of the preset first network model according to the jitter loss function, and continue to perform the step of inputting the frame image sequence into the preset first network model for processing until a preset training condition is met, to obtain a trained video recognition detection model; The jitter loss function determination module comprises: determining a first difference value according to the predicted heat map and the real label of a current frame; determining a modulation function according to the first difference value, the predicted heat map and the real label; the modulation function is provided with an upper limit, and the upper limit is a preset weight threshold; the weight threshold is used to enhance the jitter loss; determining a pixel loss function according to the first difference value; determining the jitter loss function according to the pixel loss function and the modulation function.
9. A video recognition detection apparatus, characterized by, The method comprises: a frame image sequence acquisition module, configured to obtain a frame image sequence to be processed; a target heat map acquisition module, configured to input the frame image sequence to be processed into a trained video recognition detection model for processing, and output a target heat map; The generation of the video recognition detection model comprises: determining a jitter loss function according to a predicted heat map and a real label; wherein the jitter loss function is used to eliminate the jitter effect between adjacent frame images; The determining of the jitter loss function according to the predicted heat map and the real label comprises: determining a first difference value according to the predicted heat map and the real label of a current frame; According to the first difference value, the predicted heat map and the real label, a modulation function is determined; the modulation function is provided with an upper limit, and the upper limit is a preset weight threshold; the weight threshold is used to enhance a jitter loss; According to the first difference value, a pixel loss function is determined; According to the pixel loss function and the modulation function, a jitter loss function is determined.
10. A computer device, comprising: The computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the method according to any one of claims 1-6 or the steps of the method according to claim 7 when executing the computer program.
11. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method according to any one of claims 1-6 or the steps of the method according to claim 7.
Citation Information
Patent Citations
A method and apparatus for generating a human key point detection model
CN109508681A
Neural network training method, computing device and storage medium
CN110059605A