Training method and training device

By introducing time and space constraint terms in neural network model training, the problem of time domain discontinuity and low accuracy of neural network model when identifying moving biological poses is solved, and high accuracy recognition under occlusion and blurred images is achieved.

CN114359965BActive Publication Date: 2025-07-22BEIJING CHAOWEIJING BIOLOGICAL TECH CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111680419.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-07-22
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

When the existing neural network model recognizes moving biological poses, there is a problem that the recognition results are not smooth and continuous in the time domain and have low accuracy in the case of blurred image or occlusion of key points.

Method used

In the training process of neural network model, the time constraint term and spatial constraint term are introduced, and the position error of the key points is obtained through the tracking method to determine the time constraint term, and the spatial constraint term is determined based on the position difference value of the key points in the same image frame, and the loss function is constructed for training.

Benefits of technology

Improves the recognition accuracy of neural network models when processing occlusion and blurred images, and ensures the continuity of recognition results in the time domain.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359965B_ABST
    Figure CN114359965B_ABST
Patent Text Reader

Abstract

The present application provides a training method and a training device. The method includes: when training a neural network model, adding a time constraint term and a space constraint term to a loss function. The time constraint term is used to constrain the positions of key points in the posture of the moving organism between adjacent image frames in the image sequence, and the space constraint term is used to define the positions of key points in the posture of the moving organism in the same image frame. The neural network model trained according to the above method can ensure a high accuracy when processing occluded and blurred images, and can also ensure the continuity of the recognition results in the time domain.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a training method and a training device. Background Art

[0002] Pose recognition refers to using a neural network model to recognize and / or extract the actions and / or key points of a moving creature in an image or video.

[0003] In the prior art, when recognizing key points in an image sequence recording the behavior of a moving creature, it usually relies on extracting and matching features in a single image or a single image frame in a moving creature video. The recognition results obtained by this single recognition method are not smooth and continuous enough in the time domain. Additionally, at the same time, in the case where the image to be recognized is blurred or the key points are occluded, the key points in the image may not be correctly recognized, resulting in a low recognition accuracy. Summary of the Invention

[0004] In view of this, embodiments of this application provide a training method and a training device to improve the accuracy and recognition efficiency of a neural network model in performing moving creature pose recognition.

[0005] In a first aspect, a training method is provided. The method includes: obtaining a training sample, where the training sample is an image sequence recording the behavior of a moving creature; inputting the training sample into a neural network model to obtain a recognition result of the pose of the moving creature; training the neural network model using a loss function according to the recognition result of the pose of the moving creature; where the loss function includes a time constraint term and a space constraint term, the time constraint term is used to constrain the positions of key points in the pose of the moving creature between adjacent image frames in the image sequence, and the space constraint term is used to define the positions of key points in the pose of the moving creature in the same image frame.

[0006] Optionally, the training method further includes: determining the time constraint term according to the error between the position of the key point obtained by using a tracking method and the position of the key point in the recognition result.

[0007] Optionally, determining the time constraint term according to the error between the positions of the key points obtained by using the tracking method and the positions of the key points in the recognition result includes: using the first image frame among the m images in the training sample as the initial frame, and performing forward tracking by using the recognition result of the initial frame to obtain a first forward tracking result, where the first forward tracking result includes the tracking positions of the key points in the m-th image frame; determining a first difference between the first forward tracking result and the recognition result of the m-th image frame; using the m-th image frame among the m images as the termination frame, and performing backward tracking by using the recognition result of the termination frame to obtain a first backward tracking result, where the first backward tracking result includes the tracking positions of the key points in the first image frame; determining a second difference between the first backward tracking result and the recognition result of the first image frame; when both the first difference and the second difference are less than or equal to a preset threshold, determining the time constraint term to be 0; when the first difference and / or the second difference is greater than the preset threshold, determining the time constraint term according to the first forward tracking result and / or the first backward tracking result; where m is a positive integer greater than or equal to 2.

[0008] Optionally, determining the time constraint term according to the first forward tracking result and / or the first backward tracking result includes: performing backward tracking by using the first forward tracking result to obtain a second backward tracking result; determining the difference between the second backward tracking result and the recognition result of the first image frame as the time constraint term; or, performing forward tracking by using the first backward tracking result to obtain a second forward tracking result; determining the difference between the second forward tracking result and the recognition result of the m-th image frame as the time constraint term.

[0009] Optionally, before training the neural network model, the training method further includes: determining the differences between the positions of the multiple key points according to the positions of the multiple key points in the recognition result; and determining the spatial constraint term according to the differences.

[0010] Optionally, determining the differences between the positions of the multiple key points according to the positions of the multiple key points in the recognition result includes: determining the distance between two key points in the same image in the training sample; when the distance is within a preset range, determining the spatial constraint term to be 0; when the distance is not within the preset range, determining the spatial constraint term according to the distance.

[0011] Optionally, determining the spatial constraint term according to the distance includes: determining the spatial constraint term to be e d , where d represents the distance.

[0012] Optionally, the preset range is determined according to the mean and variance of the distance.

[0013] Optionally, the loss function further includes an error constraint term, and the error constraint term is used to constrain the error of the key points in the posture of the moving organism in the recognition result and the annotation result.

[0014] Optionally, the error loss term is a mean square error loss term.

[0015] Optionally, training the neural network model using the loss function includes:

[0016] Training the neural network model using the gradient descent method according to the loss function.

[0017] Optionally, the neural network model includes an HRNet network.

[0018] In a second aspect, a training device is provided, and the training device includes: an acquisition module, configured to acquire a training sample, where the training sample is an image sequence recording the behavior of a moving organism; an input module, configured to input the training sample into a neural network model to obtain a recognition result of the posture of the moving organism; a training module, configured to train the neural network model using a loss function according to the recognition result of the posture of the moving organism; where the loss function includes a time constraint term and a space constraint term, the time constraint term is used to constrain the positions of the key points in the posture of the moving organism between adjacent image frames in the image sequence, and the space constraint term is used to define the relative positions of multiple key points in the posture of the moving organism in the same image frame.

[0019] Optionally, before training the neural network model, the training device further includes: a first determination module, configured to determine the time constraint term according to the error between the position of the key point obtained by using a tracking method and the position of the key point in the recognition result.

[0020] Optionally, the first determination module is configured to: use the first image frame among the m images in the training sample as the initial frame, perform forward tracking using the recognition result of the initial frame to obtain a first forward tracking result, where the first forward tracking result includes the tracking positions of key points in the m-th image frame; determine a first difference between the first forward tracking result and the recognition result of the m-th image frame; use the m-th image frame among the m images as the termination frame, perform backward tracking using the recognition result of the termination frame to obtain a first backward tracking result, where the first backward tracking result includes the tracking positions of key points in the first image frame; determine a second difference between the first backward tracking result and the recognition result of the first image frame; when both the first difference and the second difference are less than or equal to a preset threshold, determine that the time constraint term is 0; when the first difference and / or the second difference is greater than the preset threshold, determine the time constraint term according to the first forward tracking result and / or the first backward tracking result; where m is a positive integer greater than or equal to 2.

[0021] Optionally, determining the time constraint term according to the first forward tracking result and / or the first backward tracking result includes: performing backward tracking using the first forward tracking result to obtain a second backward tracking result; determining the difference between the second backward tracking result and the recognition result of the first image frame as the time constraint term; or, performing forward tracking using the first backward tracking result to obtain a second forward tracking result; determining the difference between the second forward tracking result and the recognition result of the m-th image frame as the time constraint term.

[0022] Optionally, the training device further includes: a second determination module, configured to determine a difference between positions of multiple key points according to the positions of the multiple key points in the recognition result; and determine the spatial constraint term according to the difference.

[0023] Optionally, the second determination module is configured to: determine the distance between any two key points in the same image in the training sample; when the distance is within a preset range, determine that the spatial constraint term is 0; when the distance is not within the preset range, determine the spatial constraint term according to the distance.

[0024] Optionally, the second determination module is configured to: determine that the spatial constraint term is e d , where d represents the distance.

[0025] Optionally, the preset range is determined according to the mean and variance of the distance.

[0026] Optionally, the loss function further includes an error constraint term, which is used to constrain the error of the key points in the posture of the moving organism in the recognition result and the annotation result.

[0027] Optionally, the error loss term is a mean square error loss term.

[0028] Optionally, the training module is configured to: train the neural network model using the gradient descent method according to the loss function.

[0029] Optionally, the neural network model includes an HRNet network.

[0030] By introducing time and space constraints in the training process of the neural network model in this application, the neural network model can ensure a high accuracy when processing occluded and blurred images, and can also ensure the continuity of the recognition result in the time domain. Description of the Drawings

[0031] Figure 1 It is a schematic flowchart of a training method provided by an embodiment of this application.

[0032] Figure 2 It is a schematic flowchart of a method for determining a time constraint term provided by an embodiment of this application.

[0033] Figure 3 It is a schematic flowchart of a method for determining a space constraint term provided by an embodiment of this application.

[0034] Figure 4 It is a schematic flowchart of a method for determining an error constraint term provided by an embodiment of this application.

[0035] Figure 5 It is a schematic block diagram of a training device provided by an embodiment of this application.

[0036] Figure 6 It is a schematic block diagram of a training device provided by another embodiment of this application.

[0037] Figure 7 It is a schematic block diagram of an application scenario of an embodiment of this application. Detailed Embodiments

[0038] The methods and devices in the embodiments of this application can be applied to various scenarios of posture recognition of moving organisms in an image sequence. The image sequence can be multiple image frames in a video. The multiple image frames can be consecutive multiple image frames in a video. The image sequence can also be multiple images of an animal collected by an image acquisition device such as a camera. The moving organism can be an animal. The animal can be, for example, a rodent, such as a mouse.

[0039] To facilitate the understanding of the embodiments of the present application, first, taking the pose recognition of animals as an example, the background of the present application will be described in detail with examples.

[0040] The behavior of biological neurons is closely related to the activities of animals. Usually, the pose changes of animals will cause corresponding changes in neurons. Therefore, exploring the connection and interaction methods of the complex network composed of neurons under specific behaviors is very important for the fields of neuroscience and medicine. Generally, quantitative analysis methods are adopted in this field, that is, by obtaining the pose information of animals and the behavior of neurons, their corresponding relationships are determined.

[0041] The behavior of animal neurons can be obtained, for example, by using methods such as ray scanning and miniaturized multiphoton microscopes.

[0042] There are various methods for obtaining the pose information of animals. For example, the pose information of animals can be obtained by manually annotating key points in an image sequence. However, in the face of a large amount of data, the efficiency of manual processing is low and it is easy to make mistakes, and it is impossible to guarantee the accuracy of the obtained pose information.

[0043] For another example, markers (such as displacement or acceleration sensors) can also be set at the key points of the animal's body, and the pose changes of the animal can be determined according to the changes in information such as the positions of the markers. However, for rodents, due to their small size, setting markers will interfere with their natural behaviors, resulting in a decrease in the accuracy of the collected data.

[0044] For still another example, an image acquisition device such as a depth camera can also be used to locate animals in space to obtain their pose information. However, this method is sensitive to imaging conditions and scene changes and is not applicable to all occasions.

[0045] With the development of the field of artificial intelligence, the animal pose recognition method based on neural networks is gradually replacing traditional technologies. However, currently, neural network models usually do not consider the motion law of key points of moving organisms in an image sequence over time and / or the relationship between different key points on the same image frame during training. These neural network models have the following problems in the pose recognition process:

[0046] When recognizing the animal postures in an image sequence, a neural network model usually recognizes based on each frame image itself. For example, the image sequence to be recognized includes a first frame image and a second frame image in chronological order. The neural network model recognizes the animal posture in the first frame image according to the image of the first frame, and obtains a first posture recognition result corresponding to the first frame image. The neural network model recognizes the animal posture in the second frame image according to the image of the second frame, and obtains a second posture recognition result corresponding to the second frame image. Using the above method of directly recognizing the animal posture using the current frame image, the obtained recognition result is not smooth enough in the time domain. In addition, when there are blurred or occluded images in the collected image sequence, taking rodents as an example, when the tail of a mouse is curled or occluded, the accuracy of the key point position information output by the neural network model is relatively low.

[0047] In addition, existing neural network models usually construct a loss function based on the error between the recognition result and the manually annotated result, and are trained using the backpropagation algorithm. This neural network model does not consider the continuous change of key points in the time domain and the influence of the positional relationship of each key point in space during training, resulting in a problem of relatively low accuracy when performing motion biological posture recognition. On the other hand, using the error between the above recognition result and the manually annotated result to construct a loss function to train the neural network model usually makes the initial training process slower.

[0048] In view of the above problems, embodiments of the present application provide a training method and a training device. The method provided by the embodiments of the present application introduces time constraints and space constraints during the training process of the neural network model, so that the neural network model can have a relatively high accuracy when processing occluded and blurred images, and effectively suppresses the jitter phenomenon of the recognition result of the neural network model in the time domain.

[0049] The following combines Figures 1-4 , and details the training method provided by the embodiments of the present application. Figure 1 is a schematic flowchart of the training method provided by the embodiments of the present application. Figure 1 The training method shown may include steps S11-S13.

[0050] Step S11, obtain training samples.

[0051] In an embodiment of the present application, the training samples may include an image sequence recording the behavior of a moving organism and a marking result. It can be understood that the marking result may include the position information of a preset number of key points on the body of the moving organism. For example, the key points may be body joint points and key parts. Taking animals as an example, the key points may be the joint points on the limbs of the animal, as well as the tail, eyes, nose, ears, etc. The position information may be the coordinate information of the key points.

[0052] The embodiments of the present application do not limit the method for obtaining the pre-annotated results. For example, the frame-by-frame annotation of the image frames in the image sequence can be performed using the manual annotation method. As a possible implementation, other methods with higher confidence can also be used for annotation.

[0053] There are many ways to obtain the training samples, and the embodiments of the present application do not limit this either. For example, as an implementation, an image sequence directly obtained by an image acquisition device (such as a camera, a webcam, a medical imaging device, a lidar, etc.) can be used, and the image sequence may include multiple images of a moving organism arranged in chronological order. Or, for example, the training samples can be obtained from a server (such as a local server or a cloud server, etc.). Or, the training samples can also be obtained on the network or other content platforms. For example, open-source training datasets such as the MSCOCO dataset, the MPII dataset, and the POSETTRACK dataset can be used; or, it can also be an image sequence pre-stored locally.

[0054] Step S12: Input the training samples obtained in the previous step S11 into the neural network model to obtain the recognition result of the posture of the moving organism.

[0055] The embodiments of the present application do not specifically limit the neural network model, and any neural network model capable of implementing the posture recognition described in the present application can be used. For example, the neural network model can be a 2D convolutional neural network such as VGG, ResNet, HRNet, etc. Optionally, HRNet (High Resolution Network) can maintain a high resolution throughout the feature extraction process and can perform cross-fusion of features with different resolutions during the feature extraction process. It is particularly suitable for application scenarios such as semantic segmentation, human posture, image classification, facial landmark detection, and general object recognition.

[0056] Among them, the recognition result may include the position information of a preset number of body key points of the moving organism recognized by the neural network model (which can also be simply referred to as the recognition position).

[0057] Step S13: Train the neural network model using the loss function according to the recognition result in step S12.

[0058] In some embodiments, the loss function may include a time constraint term L temporal and / or a space constraint term L spatical .

[0059] The following will separately describe in detail the determination methods of each constraint term in combination with the attached Figures 2-3 figures.

[0060] Refer to Figure 2 ,Figure 2 A method for determining time constraint items is shown.

[0061] Time constraint item L temporal can be used to constrain the positions of key points in the poses of moving organisms between adjacent image frames in an image sequence. In some embodiments, the time constraint item L temporal can be determined according to the error between the position information of the key points obtained by using a tracking method and the position information of the key nodes in the recognition result.

[0062] In the training method provided by the embodiments of the present application, the tracking method can be an unsupervised tracking method. The embodiments of the present application do not make specific limitations on the tracking method. The tracking method can be, for example, an object tracking algorithm using a regression network, an object tracking algorithm, or an optical flow method. The optical flow method can be, for example, the Lucas-Kanade optical flow method, etc.

[0063] Figure 2 The method shown may include steps S1311-S1316.

[0064] Step S1311: Use the first image frame in the m images in the training sample as the initial frame, and perform forward tracking using the recognition result of the initial frame to obtain a first forward tracking result, where the first forward tracking result includes the tracking positions of the key points in the mth image frame. Here, m is a positive integer greater than or equal to 2.

[0065] Optionally, before step S1311, Figure 2 the method shown may further include: selecting m images from the training sample.

[0066] The m images are any m images in the training sample. The m images can be consecutive m images in the training sample. It can be understood that the m images can also be all the images in the training sample.

[0067] Step S1312: Determine a first difference between the first forward tracking result and the recognition result of the mth image frame. In other words, the first difference can be the difference between the tracking position and the recognition position of the same key point in the mth image frame.

[0068] For ease of description, hereinafter, the set composed of m images will be denoted as I 1,i (i = 1, 2,..., m), and the recognition result of the set I 1,i will be denoted as where ω is the number of key points in each image frame.

[0069] Embodiments of the present application can use the first frame in the m images as the initial frame and use the recognition result of the initial frame Perform forward tracking to obtain the first forward tracking result Determine the first forward tracking result and set I 1,i The recognition result of the m-th frame in The difference F1 between them is:

[0070]

[0071] Step S1313: Use the m-th image frame in the m images as the termination frame, and perform backward tracking using the recognition result of the termination frame to obtain the first backward tracking result. The first backward tracking result includes the tracking positions of the key points in the first image frame. It can be understood that the m-th image can also be referred to as the last image frame in the m images.

[0072] Step S1314: Determine the second difference between the first backward tracking result and the recognition result of the first image frame. In other words, this second difference can be the difference between the tracking position and the recognition position of the same key point in the first image frame.

[0073] In the embodiment of the present application, the last frame in the m images can be used as the termination frame, and the recognition result of this termination frame is used to perform backward tracking to obtain the first backward tracking result Determine the first backward tracking result and set I 1,i The recognition result of the first frame in The difference F2 between them is:

[0074]

[0075] In step S1315, when both the first difference and the second difference are less than or equal to the preset threshold, determine that the time constraint term is 0.

[0076] In step S1316, when the first difference and / or the second difference is greater than the preset threshold, determine the time constraint term according to the first forward tracking result and / or the first backward tracking result.

[0077] The preset threshold is related to the motion characteristics of the organism. It should be noted that compared with the prediction result of the neural network model, the tracking result obtained by the tracking method can ensure that the tracking position of the same key point changes smoothly in the time domain. Therefore, when the difference (such as the first difference or the second difference) is less than the preset threshold, it indicates that the recognition result is close to the tracking result, and the recognition result of the neural network model is relatively smooth in the time domain. At this time, the time constraint term can be not set. When the difference is greater than the preset threshold, it indicates that the recognition result is quite different from the tracking result. That is to say, the recognition result is relatively jittery in the time domain. At this time, the neural network model can be trained by setting the time constraint term to make the recognition result output by the neural network model smoother.

[0078] The embodiments of the present application do not specifically limit the manner of determining the time constraint term. For example, the first difference can be used as the time constraint term. For another example, the second difference can be used as the time constraint term. For still another example, the first forward tracking result can be backward tracked to obtain a second backward tracking result; the time constraint term is determined according to the difference between the second backward tracking result and the recognition result of the first image frame. For still another example, the first backward tracking result can be forward tracked to obtain a second forward tracking result; the time constraint term is determined according to the difference between the second forward tracking result and the recognition result of the first image frame.

[0079] The following combines specific examples to describe in detail the manner of determining the time constraint term.

[0080] For example, the first frame among the m images can be used as the initial frame Using the recognition result of this initial frame Perform forward tracking to obtain a first forward tracking result Then, using the first forward tracking result As the termination frame, perform backward tracking to determine a second backward tracking result Determine the time constraint term as:

[0081]

[0082] For another example, the last frame among the m images can be used as the termination frame Using the recognition result of this termination frame Perform backward tracking to obtain a first backward tracking result Then, using the first backward tracking result As the initial frame, perform forward tracking to determine a second forward tracking result Determine the time constraint term as:

[0083]

[0084] The embodiment of the present application can determine the time constraint term according to the preset threshold E1 as:

[0085]

[0086] Referring to Figure 3 , Figure 3 which shows a method for determining a spatial constraint term.

[0087] The spatial constraint term L spatical can be used to limit the positions of key points in the posture of a moving organism in the same image frame. In some embodiments, the spatial constraint term L spatical can be determined according to the differences between the positions of multiple key points in the recognition result.

[0088] The method for determining the spatial constraint term L spatical provided by an embodiment of the present application may include steps S1321 - S1322.

[0089] Step S1321, determine the distance between two key points in the same image of the training sample.

[0090] The distance between the two key points can be the distance between any two key points among all the key points in the same image. It can also be the distance between any two key points among some of the key points in the same image.

[0091] For example, assume that the image includes key point 1, key point 2, key point 3, and key point 4. The distance between two key points includes the distances between key point 1 and the other three key points respectively. Or the distance between two key points can be the distance between key point 1 and any other key point. Or, the distance between two key points can be the distances between each key point and the other three key points respectively.

[0092] Optionally, before step S1321, Figure 3 the method shown may further include: selecting p images from the training sample.

[0093] The p images are any p images in the training sample. The p images can be p consecutive images in the training sample. It can be understood that the p images can also be all the images in the training sample. Wherein, p is a positive integer greater than or equal to 2.

[0094] For convenience of description, hereinafter, the set composed of m images will be denoted as I 2,j (j = 1, 2, …, p), and the recognition result of the set I 2,j will be denoted as wherein, ω is the number of key points in each image frame.

[0095] In the embodiments of the present application, the distance between two key points in the same image of the training samples may be the distance between two key points in the same image among p images.

[0096] Step S1322: When the distance is within a preset range, determine that the spatial constraint term is 0. When the distance is not within the preset range, determine the spatial constraint term according to the distance.

[0097] This preset range is related to the motion characteristics of the organism. Taking a mouse as an example, if the two key points are a front paw of the mouse and the joint point where the front limb where the front paw is located is connected to the mouse's body, the distance between the two key points is the distance between the front paw and the joint point. When the mouse's front limb is straight, the distance between the front paw and the joint point is the longest. Assume this length is a. The shortest distance between the front paw and the joint point is 0. According to the motion characteristics of the mouse, this preset range can be set as [0, a]. Therefore, when the distance between these two key points is less than or equal to a, it can be considered that the error of the recognition result is small. When the distance between these two keys in the recognition result is greater than a, it can be considered that the accuracy of the recognition result is low. At this time, the neural network model can be trained by setting the spatial constraint term to make the recognition result output by the neural network model more accurate.

[0098] Next, a specific example is used to describe in detail the method for determining the spatial constraint term.

[0099] For example, the distance d between two key points in the recognition result of the set I 2,j can be calculated.

[0100] When the distance d is within the preset range, determine that the spatial constraint term is 0. When the distance d exceeds the preset range, determine that the spatial constraint term is e d .

[0101] Optionally, the preset range can be determined according to the distribution law of the distance d. For example, when the distance d conforms to a Gaussian distribution with a mean of μ and a variance of σ 2 , different confidence levels can be selected to determine the preset range. For example, it can be {μ ± 3σ}.

[0102] In the embodiments of the present application, the spatial constraint term can be determined according to the above preset range as:

[0103]

[0104] It should be noted that the method for determining the spatial constraint term L provided in the above steps S1321 - S1323 spaticalThe method is only an example, and it can also be determined by other means. For example, the spatial constraint term can also be determined based on the error between the distances of pairwise key points in the recognition result and the distances of pairwise key points in the corresponding annotation result. The present application does not limit this.

[0105] In some embodiments, the loss function may further include an error constraint term L MSE . In some embodiments, the error constraint term L MSE can be determined according to the error between the recognition result and the position information of the same key point in the annotation result of the training sample. Refer to Figure 4 . Taking the mean square error as an example, determining the error constraint term may include steps S1331 - S1333.

[0106] Step S1331, select n images from the training samples obtained in step S11 to form a sample set I 3,k (k = 1, 2, …, n), where n is a positive integer greater than or equal to 1.

[0107] The n images are any n images in the training samples. The n images can be n consecutive images in the training samples. It can be understood that the n images can also be all the images in the training samples.

[0108] Step S1332, determine the recognition result of the sample set I 3,k and the annotation result

[0109] Step S1333, calculate the mean square error of the foregoing recognition result and the annotation result , and determine the error loss term as:

[0110]

[0111] For the error loss term, in addition to the mean square error loss, the commonly used cross - entropy loss, 0 - 1 loss, absolute value loss, etc. in the art can also be used. The method shown in the above steps S1331 - S1333 is only an example and does not limit the protection scope of the present application.

[0112] In some embodiments, the foregoing error constraint term L MSE , time constraint term L temporal and spatial constraint term L spatical can be weighted and summed to determine the loss function. That is, the loss function L = L MSE + aL temporal + bL spatical , where a and b are hyperparameters, and their values are greater than or equal to 0.

[0113] ​In some embodiments, the foregoing training method further includes training the neural network model using a loss function to obtain a trained neural network model.

[0114] There are many ways to train a neural network model, and the embodiments of this application do not limit this. For example, the gradient descent algorithm can be used to update the parameters of the neural network model according to the foregoing loss function to make the neural network model converge and obtain a trained neural network model.

[0115] The following Figure 5 describes in detail an embodiment of the training device provided by this application. It should be understood that the description of the device embodiment corresponds to that of the foregoing method embodiment. Therefore, for parts not described in detail, reference can be made to the foregoing method embodiment.

[0116] Figure 5 is a schematic block diagram of a training device 50 provided by an embodiment of this application. It should be understood that Figure 5 the illustrated device 50 is only an example, and the device 50 of the embodiments of the present invention may further include other modules or units.

[0117] It should be understood that the device 50 can execute Figures 1-4 each step in the method, and for the sake of avoiding repetition, it will not be elaborated here.

[0118] As a possible implementation, the device includes:

[0119] An acquisition module 51, configured to acquire training samples.

[0120] Among them, the training samples and the acquisition method thereof may be the same as the steps of S11 in the foregoing method, and will not be elaborated here.

[0121] An input module 52, configured to input the training samples into a neural network model to obtain an identification result of the posture of the moving organism.

[0122] A training module 53, configured to train the neural network model according to the identification result using a loss function.

[0123] Optionally, before training the neural network model, the training device further includes: a first determination module, configured to determine the time constraint term according to the error between the position of the key points obtained by using the tracking method and the position of the key points in the identification result.

[0124] Optionally, the first determination module is configured to: use the first image frame among the m images in the training sample as the initial frame, perform forward tracking using the recognition result of the initial frame to obtain a first forward tracking result, where the first forward tracking result includes the tracking positions of the key points in the m-th image frame; determine a first difference between the first forward tracking result and the recognition result of the m-th image frame; use the m-th image frame among the m images as the termination frame, perform backward tracking using the recognition result of the termination frame to obtain a first backward tracking result, where the first backward tracking result includes the tracking positions of the key points in the first image frame; determine a second difference between the first backward tracking result and the recognition result of the first image frame; when both the first difference and the second difference are less than or equal to a preset threshold, determine the time constraint term to be 0; when the first difference and / or the second difference is greater than the preset threshold, determine the time constraint term according to the first forward tracking result and / or the first backward tracking result; where m is a positive integer greater than or equal to 2.

[0125] Optionally, the determining the time constraint term according to the first forward tracking result and / or the first backward tracking result includes: performing backward tracking using the first forward tracking result to obtain a second backward tracking result; determining the difference between the second backward tracking result and the recognition result of the first image frame as the time constraint term; or, performing forward tracking using the first backward tracking result to obtain a second forward tracking result; determining the difference between the second forward tracking result and the recognition result of the m-th image frame as the time constraint term.

[0126] Optionally, the training device further includes: a second determination module, configured to determine a difference between the positions of the multiple key points according to the positions of the multiple key points in the recognition result; and determine the spatial constraint term according to the difference.

[0127] Optionally, the second determination module is configured to: determine the distance between any two key points in the same image in the training sample; when the distance is within a preset range, determine the spatial constraint term to be 0; when the distance is not within the preset range, determine the spatial constraint term according to the distance.

[0128] Optionally, the second determination module is configured to: determine the spatial constraint term to be e d , where d represents the distance.

[0129] Optionally, the preset range is determined according to the mean and variance of the distance.

[0130] Optionally, the loss function further includes an error constraint term for constraining the error of key points in the posture of the moving organism in the recognition result and the annotation result.

[0131] Optionally, the error loss term is a mean square error loss term.

[0132] Optionally, the training module is configured to: train the neural network model according to the loss function using the gradient descent method.

[0133] Optionally, the neural network model includes an HRNet network.

[0134] Optionally, the loss function includes at least one of a time constraint term \(L_{t}\) temporal , a space constraint term \(L_{s}\) spatical and an error constraint term \(L_{e}\). MSE Among them, the time constraint term \(L_{t}\) temporal , the space constraint term \(L_{s}\) spatical and the error constraint term \(L_{e}\) MSE can be determined according to the method described above. Figures 2-4 shown.

[0135] Optionally, there are many ways to train the neural network model, which are not limited in the embodiments of the present application. For example, the gradient descent algorithm can be used to update the parameters of the neural network model according to the foregoing loss function to make the neural network model converge and obtain the trained neural network model.

[0136] It should be understood that the device 50 for training the neural network model is embodied in the form of a functional module here. The term "module" here can be implemented in software and / or hardware forms, which are not specifically limited. For example, the "module" can be a software program, a hardware circuit or a combination of both to implement the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merging logic circuit and / or other suitable components to support the described functions.

[0137] As an example, the device 50 for training the neural network model provided in the embodiments of the present invention may be a processor or a chip for executing the method described in the embodiments of the present invention.

[0138] Figure 6 is a schematic block diagram of a training device 60 provided in another embodiment of the present application. Figure 6The device 60 shown includes a memory 61, a processor 62, a communication interface 63, and a bus 64. Among them, the memory 61, the processor 62, and the communication interface 63 are communicatively connected to each other through the bus 64.

[0139] The memory 61 can be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 61 can store a program. When the program stored in the memory 61 is executed by the processor 62, the processor 62 is used to execute each step of the training method provided by the embodiments of the present invention. For example, it can execute Figures 1-4 each step of the illustrated embodiment.

[0140] The processor 62 can be a general-purpose central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the training method of the method embodiment of the present invention.

[0141] The processor 62 can also be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the training method provided by the embodiments of the present invention can be completed by the integrated logic circuit in the hardware of the processor 62 or the instructions in software form.

[0142] The above-mentioned processor 62 can also be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0143] The steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 61, and the processor 62 reads the information in the memory 61 and combines its hardware to complete the functions required to be executed by the units included in the device for attitude recognition in the embodiments of the present invention, or executes the training method of the method embodiments of the present invention. For example, it can execute Figures 1-4 each step / function of the illustrated embodiment.

[0144] The communication interface 63 can use, but is not limited to, a transceiver device such as a transceiver to implement the communication between the device 60 and other devices or communication networks.

[0145] The bus 64 can include a path for transmitting information between various components of the device 60 (for example, the memory 61, the processor 62, the communication interface 63).

[0146] It should be understood that the device 60 shown in the embodiments of the present invention can be a processor or a chip for executing the method described in the embodiments of the present invention.

[0147] It should be understood that the processor in the embodiments of the present invention can be a central processing unit (CPU), and this processor can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0148] Next, in combination with Figure 7 the application scenario, the specific application of the embodiments of the present application will be introduced. It should be noted that the following description of Figure 7 is only an example and not a limitation. The method in the embodiments of the present application is not limited to this and can also be applied to other attitude recognition scenarios.

[0149] Figure 7 The application scenario in

[0150] Among them, the image acquisition device 71 can be used to acquire an image sequence of a moving organism. The image processing device 72 can be integrated in an electronic device, which can be a server or a terminal or other devices, and the embodiments of the present application do not limit this. For example, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, cloud storage, cloud communication, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a computer, and an intelligent Internet of Things device, etc. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.

[0151] A neural network model can be deployed in the image processing device 72, which can be used to identify the image by using the neural network model after the image sequence acquired by the above image acquisition device 71, and obtain the position information of the key points in the image to be processed. Among them, the position information of the key points can include, for example, the position coordinate information of the body joints, torso or facial features of the moving organism, etc.

[0152] The above electronic device can also use the image acquisition device 71 to acquire training samples, and train the neural network model by using a loss function according to the recognition result of the training samples and the manually marked result. The image processing device 72 can also identify the image to be processed through the trained neural network model, so as to achieve the purpose of accurately identifying the image.

[0153] The above-described embodiments are only a part of the embodiments of the present application, rather than all embodiments. The description order of the above embodiments does not limit the preferred order of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0154] It should be understood that in the embodiments of the present application, "B corresponding to A" means that B is associated with A, and B can be determined according to A. However, it should also be understood that determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information.

[0155] It should be understood that the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects.

[0156] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0157] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be read by a computer or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital video disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)), etc.

[0158] The above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A training method, characterized in that, Including: Obtain training samples, where the training samples are image sequences recording the behaviors of a moving organism; Input the training samples into a neural network model to obtain the recognition result of the posture of the moving organism; Train the neural network model according to the recognition result of the posture of the moving organism by using a loss function; Wherein, the loss function includes a time constraint term and a space constraint term, the time constraint term is used to constrain the positions of key points in the posture of the moving organism between adjacent image frames in the image sequence, and the space constraint term is used to define the positions of key points in the posture of the moving organism in the same image frame; Before training the neural network model, the training method further includes: Determine the time constraint term according to the error between the position of the key point in the first image frame of the m images in the training samples obtained by using a tracking method and the position of the key point in the recognition result, and the error between the position of the key point in the mth image frame and the position of the key point in the recognition result, where m is a positive integer greater than or equal to 2.

2. The training method according to claim 1, wherein The determining the time constraint term according to the error between the position of the key point in the first image frame of the m images in the training samples obtained by using a tracking method and the position of the key point in the recognition result, and the error between the position of the key point in the mth image frame and the position of the key point in the recognition result, includes: Take the first image frame of the m images in the training samples as the initial frame, perform forward tracking by using the recognition result of the initial frame to obtain a first forward tracking result, and the first forward tracking result includes the tracking positions of key points in the mth image frame; Determine a first difference between the first forward tracking result and the recognition result of the mth image frame; Take the mth image frame of the m images as the termination frame, perform backward tracking by using the recognition result of the termination frame to obtain a first backward tracking result, and the first backward tracking result includes the tracking positions of key points in the first image frame; Determine a second difference between the first backward tracking result and the recognition result of the first image frame; When both the first difference and the second difference are less than or equal to a preset threshold, determine the time constraint term to be 0; When the first difference and / or the second difference is greater than the preset threshold, determine the time constraint term according to the first forward tracking result and / or the first backward tracking result.

3. The training method according to claim 2, wherein The determining the time constraint term according to the first forward tracking result and / or the first backward tracking result includes: Perform backward tracking by using the first forward tracking result to obtain a second backward tracking result; Determine the difference between the second backward tracking result and the recognition result of the first image frame as the time constraint term; or, Perform forward tracking by using the first backward tracking result to obtain a second forward tracking result; Determine the difference between the second forward tracking result and the recognition result of the mth image frame as the time constraint term.

4. The training method according to claim 1, characterized in that Before training the neural network model, the training method further includes: Determine the differences between the positions of the multiple key points in the recognition result. Determine the spatial constraint term according to the differences.

5. The training method according to claim 4, characterized in that The step of determining the differences between the positions of the multiple key points according to the positions of the multiple key points in the recognition result includes: Determine the distance between two key points in the same image of the training sample. When the distance is within a preset range, determine the spatial constraint term to be 0. When the distance is not within the preset range, determine the spatial constraint term according to the distance.

6. The training method according to claim 5, characterized in that The step of determining the spatial constraint term according to the distance includes: Determine that the spatial constraint term is e d , where d represents the distance.

7. The training method according to claim 5, wherein The preset range is determined according to the mean and variance of the distance.

8. The training method according to claim 1, wherein The loss function further includes an error constraint term, which is used to constrain the error of the key points in the posture of the moving organism between the recognition result and the annotation result.

9. The training method according to claim 8, wherein The error loss term is the mean square error loss term.

10. The training method according to claim 1, wherein The step of training the neural network model using the loss function includes: Train the neural network model using the gradient descent method according to the loss function.

11. The training method according to any one of claims 1-10, characterized in that, The neural network model includes the HRNet network.

12. A training device, characterized in that, It includes: An acquisition module, configured to acquire a training sample, where the training sample is an image sequence recording the behavior of a moving organism. An input module, configured to input the training sample into a neural network model to obtain a recognition result of the posture of the moving organism. A training module, configured to train the neural network model using a loss function according to the recognition result of the posture of the moving organism. Wherein, the loss function includes a time constraint term and a spatial constraint term. The time constraint term is used to constrain the positions of the key points in the posture of the moving organism between adjacent image frames in the image sequence, and the spatial constraint term is used to define the relative positions of multiple key points in the posture of the moving organism in the same image frame. Before training the neural network model, the training device further includes: A first determination module, configured to determine the time constraint term according to the error between the position of the key point in the first image frame of the m images in the training sample obtained by using a tracking method and the position of the key point in the recognition result, and the error between the position of the key point in the mth image frame and the position of the key point in the recognition result, where m is a positive integer greater than or equal to 2.

13. The training device according to claim 12, wherein The first determination module is configured to: Take the first image frame of the m images in the training sample as the initial frame, and perform forward tracking using the recognition result of the initial frame to obtain a first forward tracking result, where the first forward tracking result includes the tracked positions of the key points in the mth image frame. Determine a first difference between the first forward tracking result and the recognition result of the mth image frame. Take the mth image frame of the m images as the termination frame, and perform backward tracking using the recognition result of the termination frame to obtain a first backward tracking result, where the first backward tracking result includes the tracked positions of the key points in the first image frame. Determine a second difference between the first backward tracking result and the recognition result of the first image frame. When both the first difference and the second difference are less than or equal to a preset threshold, determine that the time constraint term is 0; When the first difference and / or the second difference is greater than the preset threshold, determine the time constraint term according to the first forward tracking result and / or the first backward tracking result.

14. The training device according to claim 13, characterized in that, The determining the time constraint term according to the first forward tracking result and / or the first backward tracking result includes: Performing backward tracking using the first forward tracking result to obtain a second backward tracking result; Determining the difference between the second backward tracking result and the recognition result of the first image frame as the time constraint term; or, Performing forward tracking using the first backward tracking result to obtain a second forward tracking result; Determining the difference between the second forward tracking result and the recognition result of the m-th image frame as the time constraint term.

15. The training device according to claim 12, wherein The training device further includes: A second determination module, configured to determine the difference between the positions of multiple key points in the recognition result; Determine the space constraint term according to the difference.

16. The training device according to claim 15, wherein The second determination module is configured to: Determine the distance between any two key points in the same image of the training sample; When the distance is within a preset range, determine that the space constraint term is 0; When the distance is not within the preset range, determine the space constraint term according to the distance.

17. The training device according to claim 16, characterized in that, The second determination module is configured to: Determine that the spatial constraint term is e d , where d represents the distance.

18. The training device according to claim 16, wherein The preset range is determined according to the mean and variance of the distance.

19. The training device according to claim 12, wherein The loss function further includes an error constraint term, and the error constraint term is used to constrain the error of the key points in the posture of the moving organism in the recognition result and the annotation result.

20. The training device according to claim 19, wherein The error loss term is a mean square error loss term.

21. The training device according to claim 12, wherein The training module is configured to: Train the neural network model using the gradient descent method according to the loss function.

22. The training device according to any one of claims 12-21, characterized in that, The neural network model includes an HRNet network.

Citation Information

Patent Citations

  • Localisation, mapping and network training

    CN111902826A