Image recognition model training method and device, computer device, and storage medium
By changing the architectural orientation of sample objects during image recognition model training, sample images and labels under different orientations are obtained, solving the problem of inaccurate recognition during image recognition model training and achieving accurate and high-precision recognition results under different orientations.
Patent Information
- Application Number
- CN202310896943.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-20
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-07-20
AI Technical Summary
Existing image recognition models are prone to inaccurate recognition results for specific locations in images during the training process, resulting in poor practicality.
By acquiring sample images from multiple frames, changing the architectural orientation of the sample object in each frame, and obtaining sample orientation images under different architectural orientations, and obtaining training labels corresponding to the symptom locations of the sample object, the image recognition model is trained using these images and labels to enhance image features and ensure recognition accuracy under different architectural orientations.
This improved the accuracy of the trained image recognition model under different architectural orientations, enhanced the model's practicality, ensured accurate recognition even with different object architectural orientations, and confirmed the achievement of the preset recognition accuracy through test images, thus improving the accuracy and practicality of the recognition model.
Smart Images

Figure CN116895000B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image recognition, and in particular to a training method and device of an image recognition model, a computer device and a storage medium. BACKGROUND
[0002] With the development of science and technology, the identification of specific positions in images is gradually transformed into intelligence.
[0003] Most image recognition methods are to recognize images through trained image recognition models, but the training process of the image recognition model usually adopts the method of inputting sample images into the image recognition model. Since the input is a sample image, it is easy to cause inaccurate recognition results of specific positions in the image.
[0004] Therefore, the trained image recognition model in the prior art has poor practicability. SUMMARY
[0005] Therefore, it is necessary to provide a training method and device of an image recognition model, a computer device and a storage medium, which can improve the practicability of the trained image recognition model.
[0006] In a first aspect, the present application provides a training method of an image recognition model. The method comprises:
[0007] Obtaining a sample image of a sample object, the sample image comprising a plurality of images;
[0008] Obtaining sample directional images under different architecture directions by changing the architecture direction of the sample object in each image;
[0009] Obtaining training labels corresponding to the symptom position of the sample object in each image of each sample directional image;
[0010] Training the image recognition model through each sample directional image and the corresponding training label to obtain a trained image recognition model under each architecture direction.
[0011] In one embodiment, obtaining training labels corresponding to the symptom position of the sample object in each image of each sample directional image comprises:
[0012] For each image, determining the skeleton of the sample object in the image;
[0013] Determining a skeleton line segment of the skeleton having a symptom;
[0014] Determining an image region containing the skeleton line segment in the image as a training label.
[0015] In one of the embodiments, the image recognition model is trained by the sample direction images and the corresponding training labels, including:
[0016] In the process of processing the image features in the sample direction images by the image recognition model, the corresponding image features are enhanced based on the training labels, and the image recognition model is trained according to the enhanced image features.
[0017] In one of the embodiments, the method further includes:
[0018] Obtaining a test image for a test object, the test image including multiple frames of images;
[0019] For a current frame of test image in the test image, a test direction image under a different architecture direction is obtained by changing the architecture direction of the test object in the current frame of test image;
[0020] The test direction images under different architecture directions are recognized by the trained image recognition model to obtain corresponding recognition results;
[0021] The test labels of the test direction images under different architecture directions are compared with the corresponding recognition results in terms of similarity to obtain the corresponding similarity of the test direction images under different architecture directions;
[0022] The similarities under different architecture directions are integrated to obtain a cumulative similarity corresponding to the current frame of test image;
[0023] In the case where the cumulative similarities corresponding to the images in the test image, which are continuous for a first preset number of frames, all reach a preset threshold, it is determined that the trained image recognition model reaches a preset recognition accuracy.
[0024] In one of the embodiments, the test labels of the test direction images under different architecture directions are compared with the corresponding recognition results in terms of similarity to obtain the corresponding similarity of the test direction images under different architecture directions, including:
[0025] For the test direction image under each architecture direction, a skeleton of the test object in the test direction image is determined;
[0026] Based on the number of angles formed by the line segments composed of the joint nodes in the skeleton, a plurality of dimensions for constructing a vector are determined;
[0027] The joint nodes matching the corresponding test labels of the test direction image are determined, the value of each dimension in the plurality of dimensions is determined, and the value of the corresponding dimension of the matching joint nodes is enhanced to obtain a first vector;
[0028] determine a joint node corresponding to the recognition result of the test direction image, determine a value of each dimension in the plurality of dimensions, and enhance the value of the corresponding dimension of the matching joint node, to obtain a second vector;
[0029] calculate a similarity between the first vector and the second vector as a similarity corresponding to the test direction image.
[0030] In one of the embodiments, the method further comprises:
[0031] acquiring a real-time target image for the target object during movement of the target object along the guide route;
[0032] performing recognition on the current frame target image in the real-time target image through an image recognition model reaching a preset recognition accuracy, to obtain a recognition result;
[0033] in a case where the recognition result is that there is a symptom area, taking the current frame target image as a starting frame of cutting;
[0034] for a subsequent image frame of the starting frame of cutting, determining a candidate image frame in which the recognition result of the subsequent image frame is that there is no symptom area;
[0035] in a case where the recognition result of the image frame is that there is no symptom area from the candidate image frame and for a continuous second preset number of image frames, determining the corresponding candidate image frame as an ending frame of cutting;
[0036] based on the starting frame of cutting and the ending frame of cutting, cutting the real-time target image to obtain a target image segment.
[0037] In a second aspect, the application further provides a training device of an image recognition model. The device comprises:
[0038] a sample image acquisition module configured to acquire a sample image for a sample object, the sample image comprising a plurality of images;
[0039] an architecture direction changing module configured to acquire sample direction images under different architecture directions by changing an architecture direction of the sample object in each image;
[0040] a label acquisition module configured to acquire a training label corresponding to a symptom position of the sample object in each image of each sample direction image;
[0041] a model training module configured to train the image recognition model through each sample direction image and the corresponding training label, to obtain a trained image recognition model under each architecture direction.
[0042] In a third aspect, the present application also provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the following steps when executing the computer program:
[0043] obtaining a sample image for a sample object, the sample image comprising a plurality of images;
[0044] obtaining sample directional images in different architectural directions by changing the architectural direction of the sample object in each image;
[0045] obtaining training labels corresponding to the position of the disease sign of the sample object in each image of each sample directional image;
[0046] training an image recognition model through each sample directional image and the corresponding training label, and obtaining a trained image recognition model in each architectural direction.
[0047] In a fourth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the following steps: obtaining a sample image for a sample object, the sample image comprising a plurality of images;
[0048] obtaining sample directional images in different architectural directions by changing the architectural direction of the sample object in each image;
[0049] obtaining training labels corresponding to the position of the disease sign of the sample object in each image of each sample directional image;
[0050] training an image recognition model through each sample directional image and the corresponding training label, and obtaining a trained image recognition model in each architectural direction.
[0051] In a fifth aspect, the present application also provides a computer program product. The computer program product comprises a computer program, and the computer program is executed by a processor to implement the following steps:
[0052] obtaining a sample image for a sample object, the sample image comprising a plurality of images;
[0053] obtaining sample directional images in different architectural directions by changing the architectural direction of the sample object in each image;
[0054] obtaining training labels corresponding to the position of the disease sign of the sample object in each image of each sample directional image;
[0055] training an image recognition model through each sample directional image and the corresponding training label, and obtaining a trained image recognition model in each architectural direction.
[0056] The training method, device, computer device and storage medium of the image recognition model provided by the embodiments of the present application can obtain sample images including multiple frames of images for sample objects, change the architecture direction of the sample objects in each frame of image, obtain sample direction images under different architecture directions, and obtain training labels corresponding to the position of the symptom of the sample object in each frame of image of each sample direction image. The image recognition model is trained by using each sample direction image and the corresponding training label, and the trained image recognition model under each architecture direction is obtained. Compared with the problem of poor practicability of the trained image recognition model in the prior art, the embodiments of the present application can obtain the recognition result of continuous frames by obtaining sample images with continuous frames and taking the sample images as training samples, thereby avoiding the inaccurate recognition caused by recognizing one frame of image. In addition, the embodiments of the present application can change the architecture direction of the sample objects in each frame of image, obtain sample direction images under different architecture directions, take the sample direction images under different architecture directions as training samples, obtain the trained image recognition model under different architecture directions, and ensure the accuracy of the recognition result of the trained image recognition model under the condition that the architecture direction of the object in the image is different, thereby improving the practicability of the trained image recognition model. BRIEF DESCRIPTION OF DRAWINGS
[0057] Figure 1 A flowchart of the training method of the image recognition model provided in the embodiments of the present application is shown.
[0058] Figure 2 A schematic diagram of a symptom position is provided in one embodiment.
[0059] Figure 3 A schematic diagram of changing the architecture direction of a sample object in a frame of image is provided in one embodiment.
[0060] Figure 4 A flowchart of obtaining a training label is provided in one embodiment.
[0061] Figure 5 A schematic diagram of the skeleton of a sample object in a frame of image is provided in one embodiment.
[0062] Figure 6 A schematic diagram of an image region containing a skeleton line segment is provided in one embodiment.
[0063] Figure 7 A flowchart of determining that the trained image recognition model reaches the preset recognition accuracy is provided in one embodiment.
[0064] Figure 8 A flowchart of obtaining the similarity of the test direction images under different architecture directions is provided in one embodiment.
[0065] Figure 9 Fig. 1 shows an angle formed by a line segment composed of a joint in a skeleton according to an embodiment;
[0066] Figure 10 Fig. 2 shows a test label corresponding to a test direction image according to an embodiment;
[0067] Figure 11 Fig. 3 shows a recognition result corresponding to a test direction image according to an embodiment;
[0068] Figure 12 Fig. 4 shows a flowchart of obtaining a target image segment according to an embodiment;
[0069] Figure 13 Fig. 5 shows a schematic diagram of a target object moving along a guide route according to an embodiment;
[0070] Figure 14 Fig. 6 shows a structural block diagram of a training device of an image recognition model according to an embodiment;
[0071] Figure 15 Fig. 7 shows an internal structure diagram of a computer device according to an embodiment. DETAILED DESCRIPTION
[0072] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.
[0073] In the present embodiment, as shown in Figure 1 Fig. 1, a training method of an image recognition model is provided, and the present embodiment takes the method applied to a computer device as an example for illustration. It should be understood that the computer device can be a terminal or a server, but is not limited thereto. The terminal can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The method can also be applied to a system including a computer device and a server, and can be implemented through the interaction of the computer device and the server. In the present embodiment, the method includes the following steps:
[0074] S101, obtaining a sample image for a sample object, the sample image including multiple frames of images.
[0075] The sample object is an object including a symptom position. The sample object can be, but is not limited to, a human, and other movable objects including a symptom position can also be used as the sample object. The sample image is an image of the sample object during movement.
[0076] A schematic diagram of a symptom position is provided as shown in Figure 2 Figure 2 The four images shown in the figure represent four images of the same frame in different architectural directions.
[0077] In S102, sample directional images in different architectural directions are obtained by changing the architectural direction of the sample object in each frame of image.
[0078] The architectural direction can be understood as the direction of the sample object viewed from different directions at a certain moment, for example, the architectural direction can be 45 degrees, 90 degrees, 135 degrees, 180 degrees, 225 degrees, 270 degrees, 315 degrees, and 360 degrees.
[0079] In some embodiments, the architectural direction of the sample object in each frame of image is changed by adjusting the architectural direction of the sample object in each frame of image through a rotation matrix. Other methods that can change the architectural direction of the sample object in each frame of image can also be used, for example, the sample object can be photographed at different angles, or the architectural direction of the sample object in each frame of image can be directly changed. The specific implementation is not limited.
[0080] In this embodiment, for a sample image, sample directional images in different architectural directions are obtained, that is, the sample image is expanded into multiple sample directional images, which can increase the input samples of the image recognition model, and the training result of the image recognition model is more accurate.
[0081] Specifically, a schematic diagram of changing the architectural direction of the sample object in a frame of image is provided as shown in Figure 3 Figure 3 The four images shown in the figure represent four images of the same frame in different architectural directions.
[0082] In S103, training labels corresponding to the symptom position of the sample object in each frame of image of each sample directional image are obtained.
[0083] The training label is used to represent the position of the symptom position of the sample object in each frame of image. The training label can be in the form of coordinates, or other forms that can mark the symptom position of the sample object. The specific implementation is not limited.
[0084] In S104, the image recognition model is trained by using each sample directional image and the corresponding training label, and a trained image recognition model in each architectural direction is obtained.
[0085] The image recognition model is a model used to identify features in an image. For example, the image recognition model can employ a long short-term memory recurrent neural network.
[0086] The image recognition model training method provided in this embodiment acquires sample images of a sample object, including multiple frames, changes the architectural orientation of the sample object in each frame, and acquires sample orientation images under different architectural orientations. It also acquires training labels corresponding to the symptom locations of the sample object in each frame of each sample orientation image. The image recognition model is trained using these sample orientation images and corresponding training labels to obtain a trained image recognition model for each architectural orientation. Compared to traditional techniques where the trained image recognition model has poor practicality, this embodiment, by acquiring sample images with consecutive frames and using them as training samples, can obtain recognition results for consecutive frames, avoiding inaccuracies caused by recognizing only a single frame. Furthermore, this embodiment can change the architectural orientation of the sample object in each frame, acquiring sample orientation images under different architectural orientations, and using these images as training samples to obtain trained image recognition models for different architectural orientations. This ensures the accuracy of the recognition results even when the architectural orientation of the object in the image is different, improving the practicality of the trained image recognition model.
[0087] In one embodiment, training labels corresponding to the symptom locations of sample objects in each frame of each sample orientation image are obtained. A flowchart illustrating the process of obtaining training labels is provided, such as... Figure 4 As shown, it includes the following:
[0088] S401, for each frame of image, determine the skeleton of the sample object in the image.
[0089] The skeleton of a sample object in an image is formed by connecting the joints of the sample object.
[0090] In some embodiments, for each frame of an image, the skeleton of the sample object in the image is determined by using a human skeleton key point detection algorithm to determine the skeleton of the sample object in each frame of the image. Alternatively, the skeleton of the sample object in each frame of the image can be manually annotated. The specific method is not limited.
[0091] A schematic diagram of the skeleton of the sample object in a frame of an image, as shown below. Figure 5 As shown. Figure 5 The object in the middle is the sample object, and the lines connecting its body represent its skeleton. (The above...) Figure 3 The lines connecting the sample objects in four frames of images with different architectural orientations within the same frame also represent the skeleton.
[0092] S402, determine a skeleton line segment with a symptom in the skeleton.
[0093] In some embodiments, the determination of the skeleton line segment with the symptom in the skeleton can be to first determine a symptom joint in the skeleton, to take a line segment with a first preset length passing through the joint as the skeleton line segment with the symptom, or to take a line connecting two joints as the skeleton line segment with the symptom, or to directly take a second preset length of a line segment between the two joints as the skeleton line segment with the symptom, and the specific manner is not limited.
[0094] S403, determine an image region containing the skeleton line segment in the image as a training label.
[0095] The image region containing the skeleton line segment is a region within a preset range containing the skeleton line segment on the image. The preset range is usually a range not including other skeletons around the skeleton line segment.
[0096] Optionally, the coordinates of the skeleton line segment are obtained as the training label.
[0097] Specifically, the determination of the image region containing the skeleton line segment in the image can be in a manual smearing manner to determine the image region containing the skeleton line segment, and the boundary contour coordinates of the image region are taken as the training label. Accordingly, a schematic diagram of an image region containing a skeleton line segment is provided, as shown in Figure 6 , wherein A in Figure 6 represents the image region. It should be understood that the image region in Figure 6 is only one possible case and does not constitute a limitation on the image region.
[0098] In the embodiment, the training label is obtained to train the image recognition model subsequently, which can reflect the training quality of the image recognition model. The training label is obtained by marking the skeleton to ensure consistency. The sample objects have body size differences, and the training label obtained by the skeleton is more accurate and stable.
[0099] In one embodiment, the image recognition model is trained by the sample directional image and the corresponding training label, including:
[0100] In the process of processing the image features in the sample directional image by the image recognition model, the corresponding image features are enhanced based on the training label, and the image recognition model is trained according to the enhanced image features.
[0101] The manner of enhancing the corresponding image features is a weighted proportion, and the value of the weighted proportion can be set by itself.
[0102] In the embodiment, the image recognition model is trained according to the enhanced image features, so that the identified symptom positions are more accurate, and the trained image recognition model has better recognition effect.
[0103] In one embodiment, the training method of the image recognition model further includes a process of determining that the trained image recognition model reaches a preset recognition accuracy. In the embodiment, a flowchart for determining that the trained image recognition model reaches the preset recognition accuracy is provided, as shown in Figure 7
[0104] S701, obtaining a test image for a test object, the test image including multiple frames of images.
[0105] The test object is an object including a symptom position. The test object is usually a person, and is not specifically limited. Other movable objects including symptom positions can also be used as the test object. The test object can be the same as or different from the sample object, and is preferably different from the sample object. The test image is an image for the test object during movement.
[0106] S702, for a current frame of test image in the test image, obtaining a test direction image under a different architecture direction by changing the architecture direction of the test object in the current frame of test image.
[0107] Optionally, the manner of changing the architecture direction of the test object in the current frame of test image is adjusting the architecture direction of the test object in the current frame of test image by a rotation matrix. Other manners of changing the architecture direction of the test object in the current frame of test image can also be used, and are not specifically limited.
[0108] Each current frame of test image corresponds to a test direction image under a different architecture direction. The architecture direction of the test direction image needs to be consistent with the architecture direction of the sample direction image.
[0109] S703, identifying the test direction images under different architecture directions by the trained image recognition model to obtain corresponding recognition results.
[0110] The recognition result includes the symptom position in the test direction image.
[0111] S704, comparing the test labels of the test direction images under different architecture directions with the corresponding recognition results in terms of similarity to obtain the corresponding similarities of the test direction images under different architecture directions.
[0112] The test label is used to represent the position of the symptom position of the test object in the test direction image.
[0113] Specifically, the manner of obtaining the test label comprises: determining a skeleton of the test object in the test direction image, determining a skeleton line segment with a symptom in the skeleton, and determining an image region containing the skeleton line segment in the test direction image as the test label.
[0114] S705, integrating according to the similarities under different architecture directions to obtain a cumulative similarity corresponding to the current frame test image.
[0115] The integration according to the similarities under different architecture directions can adopt a manner of directly adding the similarities under different architecture directions, or a manner of weighted summation of the similarities under different architecture directions, and the specific manner is not limited.
[0116] S706, in a case where the respective cumulative similarities of the images each existing for a continuous first preset number of frames in the test image reach a preset threshold, determining that the trained image recognition model reaches a preset recognition accuracy.
[0117] The first preset number of frames can be set, and the specific number of frames is not limited. The preset threshold can be pre-set by the computer device, and the specific manner is not limited. The preset threshold and the first preset number of frames determine the preset recognition accuracy, and the preset recognition accuracy can be changed by changing at least one of the preset threshold or the first preset number of frames.
[0118] In the embodiment, the trained image recognition model is determined to reach the preset recognition accuracy through the test image, so as to ensure that the accuracy of the trained image recognition model is good, and the accuracy of the recognition result of the image recognition model reaching the preset recognition accuracy is improved.
[0119] In one embodiment, the test label of the test direction image under different architecture directions is compared with the corresponding recognition result to obtain the similarity, and a flowchart of the similarity of the test direction image under different architecture directions is shown as Figure 8 , which includes the following contents:
[0120] S801, for each test direction image under each architecture direction, determining a skeleton of the test object in the test direction image.
[0121] In some embodiments, for each test direction image under each architecture direction, the manner of determining the skeleton of the test object in the test direction image is to determine the skeleton of the test object in the test direction image through a human body skeleton key point detection algorithm, and the manner of manually labeling the skeleton of the test object in the test direction image can also be used, and the specific manner is not limited.
[0122] S802, determining a plurality of dimensions for constructing a vector based on the number of angles formed by the line segments composed of the joints in the skeleton.
[0123] For the convenience of understanding, a schematic diagram of the angle formed by the line segment composed of the joint in the skeleton is provided, as shown in Figure 9 B in Figure 9 represents the angle.
[0124] Based on the number of angles formed by the line segment composed of the joint in the skeleton, the number of dimensions used to construct the vector is determined, that is, the determination method of the dimension. The number of all angles formed by the line segment composed of the joint in the skeleton can be taken as the dimension of the constructed vector, or part of the key angles formed by the line segment composed of the joint in the skeleton can be selected, and the number of part of the key angles can be taken as the dimension of the constructed vector.
[0125] S803, determining the joint matched with the test label corresponding to the test direction image, determining the value of each dimension in the plurality of dimensions, and enhancing the value of the corresponding dimension of the matched joint to obtain a first vector.
[0126] In some embodiments, a schematic diagram of a test label corresponding to a test direction image is provided, as shown in Figure 10 The area circled in Figure 10 is the test label. It should be understood that the test label can not be one, but can be multiple.
[0127] The method for determining the joint matched with the test label corresponding to the test direction image is to screen the joint contained in the image area corresponding to the test label in the test direction image, and screen out the joint with the angle value as the matched joint.
[0128] Taking the first vector as an example, for example, the first vector is (aX1, X2, X3), wherein X1, X2 and X3 are the angle values of three key angles, X1 is the angle value of the corresponding dimension of the joint matched with the test label, and a is an enhancement coefficient. Specifically, the value of each dimension of the first vector can not be an angle value, the values of X1, X2 and X3 can be set to 1, and the value of the enhancement coefficient a is greater than 1.
[0129] S804, determining the joint matched with the recognition result corresponding to the test direction image, determining the value of each dimension in the plurality of dimensions, and enhancing the value of the corresponding dimension of the matched joint to obtain a second vector.
[0130] In a feasible implementation, a schematic diagram of a recognition result corresponding to a test direction image is provided, as shown in Figure 11 The area circled in Figure 11 is the recognition result. It should be understood that the recognition result can not be one, but can be multiple.
[0131] The way of determining the joint node matched with the recognition result corresponding to the test direction image is to screen the joint nodes contained in the image area corresponding to the recognition result on the test direction image, and screen out the joint nodes with the included angle value as the matched joint node.
[0132] Taking the second vector as an example, the second vector is (X1, bX2, X3), where X1, X2 and X3 are the included angle values of the three key included angles, X2 is the included angle value of the corresponding dimension of the joint node matched with the recognition result, and b is the enhancement coefficient. Specifically, the value of each dimension of the second vector can not be an included angle value, the values of X1, X2 and X3 can be set to 1, and the value of the enhancement coefficient b is greater than 1. The value of the enhancement coefficient b can be consistent with the value of the enhancement coefficient a, or can be inconsistent with the value of the enhancement coefficient a.
[0133] S805, calculating the similarity between the first vector and the second vector as the similarity corresponding to the test direction image.
[0134] There are many ways to calculate the similarity between the first vector and the second vector, and the calculation method of the similarity between the vectors is not limited.
[0135] In this embodiment, by comparing the similarity of the test label of the test direction image and the corresponding recognition result, if the recognition result is consistent with the test label, the similarity will be larger, which indicates that the recognition result of the trained image recognition model is more accurate. By comparing the similarity, the recognition accuracy of the trained image recognition model is determined, and the accuracy of the trained image recognition model is further ensured to be good.
[0136] In one embodiment, the training method of the image recognition model further includes the process of obtaining a target image segment. In this embodiment, a flowchart for obtaining a target image segment is provided, as shown in Figure 12 The flowchart includes the following steps:
[0137] S1201, during the movement of the target object along the guide route, obtaining a real-time target image for the target object.
[0138] The target object is an object to be identified for a disease position. The real-time target image is an image for the target object during movement. The real-time target image is obtained through a real-time transmission stream protocol.
[0139] In one possible implementation, a schematic diagram of the movement of a target object along a guide route is provided, as shown in Figure 13 The dashed line in Figure 13 represents the guide route.
[0140] S1202, for the current frame target image in the real-time target image, the current frame target image is identified through the image recognition model reaching the preset recognition accuracy, and an identification result is obtained.
[0141] The identification result includes whether there is a symptom position in the current frame target image, and in the case where there is a symptom position, the symptom position is marked in the current frame target image.
[0142] S1203, in the case where the identification result is a symptom area, the current frame target image is taken as a starting cutting frame.
[0143] The starting cutting frame is used to represent the starting point of cutting the real-time target image.
[0144] S1204, for the subsequent image frames of the starting cutting frame, the candidate image frames in which the identification result of the subsequent image frames is not a symptom area are determined.
[0145] S1205, in the case where the identification result of the image frames existing continuously for a second preset frame number from the candidate image frames is not a symptom area, the corresponding candidate image frame is determined as an ending cutting frame.
[0146] The second preset frame number can be set, and the specific frame number is not limited. The ending cutting frame is used to represent the end point of cutting the real-time target image.
[0147] S1206, based on the starting cutting frame and the ending cutting frame, the real-time target image is cut to obtain a target image segment.
[0148] In this embodiment, the real-time target image is cut to obtain a target image segment, which can identify the target object in real time, cut the effective segment as the target image segment, shorten the length of the target image segment, and improve the practicability of the target image segment.
[0149] Here, the training method of the image recognition model provided by the present application is described in detail in the form of a complete embodiment. The implementation process of the training method of the image recognition model includes:
[0150] First, the image recognition model reaching the preset recognition accuracy is obtained, and the process is as follows:
[0151] Obtain the sample image for the sample object, the sample image includes multiple image frames, and obtain the test image for the test object, the test image includes multiple image frames;
[0152] Adjust the architecture direction of the sample object in each image frame through the rotation matrix, and obtain the sample direction image under different architecture directions;
[0153] For each frame of image, determine the skeleton of the sample object in the image, determine the skeleton line segment with the symptom in the skeleton, determine the image region containing the skeleton line segment in the image as the training label;
[0154] In the process of processing the image features in the sample direction image by the image recognition model, the corresponding image features are enhanced based on the training label, and the image recognition model is trained according to the enhanced image features to obtain the trained image recognition model under each architecture direction;
[0155] For the current frame test image in the test image, the architecture direction of the test object in the current frame test image is adjusted by the rotation matrix to obtain the test direction image under different architecture directions;
[0156] The trained image recognition model is used to identify the test direction image under different architecture directions to obtain the corresponding recognition result;
[0157] For each test direction image under each architecture direction, determine the skeleton of the test object in the test direction image, determine the number of dimensions for constructing the vector based on the angle formed by the line segment composed of the joint nodes in the skeleton, determine the joint node matched with the test label of the test direction image, determine the value of each dimension in the multiple dimensions, and enhance the value of the corresponding dimension of the matched joint node to obtain the first vector, determine the joint node matched with the recognition result of the test direction image, determine the value of each dimension in the multiple dimensions, and enhance the value of the corresponding dimension of the matched joint node to obtain the second vector, and calculate the similarity between the first vector and the second vector as the similarity of the test direction image;
[0158] The similarities of the test direction images under different architecture directions are added to obtain the cumulative similarity of the current frame test image;
[0159] In the case that the cumulative similarity of each of the images in the test image with a continuous first preset number of frames reaches a preset threshold, it is determined that the trained image recognition model reaches a preset recognition accuracy.
[0160] Then, the image recognition model reaching the preset recognition accuracy is used for target image segment acquisition, and the process is as follows:
[0161] During movement of the target object along the guidance route, a real-time target image for the target object is acquired, for a current frame target image in the real-time target image, a current frame target image is recognized by an image recognition model reaching a preset recognition accuracy, a recognition result is obtained, in a case where the recognition result is that a symptom area exists, the current frame target image is taken as a starting clipping frame, for a subsequent image frame of the starting clipping frame, a candidate image frame in which the recognition result is that a symptom area does not exist is determined, in a case where the recognition result of the image frame existing for a continuous second preset number of frames from the candidate image frame is that a symptom area does not exist, the corresponding candidate image frame is determined as an ending clipping frame, and the real-time target image is clipped based on the starting clipping frame and the ending clipping frame to obtain a target image segment.
[0162] Finally, the target image segment can be transmitted to a viewing host for viewing.
[0163] The training method of the image recognition model provided in the present application can obtain the recognition result of the continuous frames by acquiring the sample image existing for a continuous frame and taking the sample image as a training sample, avoid the inaccurate recognition caused by recognizing one frame of image, and change the architecture direction of the sample object in each frame of image, acquire the sample direction image under different architecture directions, take the sample direction image under different architecture directions as a training sample, obtain the trained image recognition model under different architecture directions, so that the trained image recognition model can still ensure the accuracy of the recognition result in the case where the architecture direction of the object in the image is different, improve the practicality of the trained image recognition model, and further determine that the trained image recognition model reaches the preset recognition accuracy through the test image, so as to ensure that the accuracy of the trained image recognition model is good, improve the accuracy of the recognition result of the image recognition model reaching the preset recognition accuracy, so that the real-time target image can be recognized in real time in the process of applying the image recognition model reaching the preset recognition accuracy, has a certain real-time performance, and the effective segment is clipped as the target image segment, the length of the target image segment is shortened, and the target image segment can be transmitted to the viewing host for viewing, thereby improving the practicality of the target image segment.
[0164] It should be understood that although the steps in the flowcharts involved in the embodiments described above are shown in sequence according to the arrows, the steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, the execution of the steps is not strictly limited in sequence, and the steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the embodiments described above can include multiple steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of the steps or stages is not necessarily sequential, but can be alternately or alternately executed with at least part of other steps or steps or stages in other steps.
[0165] Based on the same inventive concept, the embodiments of the present application also provide a training device for the image recognition model involved in the training method of the image recognition model. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more image recognition model training device embodiments provided below can refer to the limitations of the image recognition model training method in the above text, which will not be repeated here.
[0166] Referring to Figure 14 , Figure 14 A structural block diagram of an image recognition model training device provided in an embodiment of the present application, the device 1400 includes: a sample image acquisition module 1401, an architecture direction changing module 1402, a label acquisition module 1403 and a model training module 1404, wherein:
[0167] The sample image acquisition module 1401 is configured to acquire a sample image of a sample object, the sample image including multiple frames of images;
[0168] The architecture direction changing module 1402 is configured to change the architecture direction of the sample object in each frame of image to obtain a sample direction image in different architecture directions;
[0169] The label acquisition module 1403 is configured to acquire a training label corresponding to a disease position of the sample object in each frame of image of each sample direction image;
[0170] The model training module 1404 is configured to train the image recognition model by using each sample direction image and the corresponding training label to obtain a trained image recognition model in each architecture direction.
[0171] The training device of the image recognition model provided in the embodiment obtains sample images including multiple frames of images for sample objects through a sample image acquisition module, changes the architecture direction of the sample objects in each frame of image through an architecture direction changing module, obtains sample direction images in different architecture directions, and obtains training labels corresponding to the disease position of the sample objects in each frame of image of each sample direction image through a label acquisition module. The model training module trains the image recognition model through each sample direction image and the corresponding training label, and obtains the trained image recognition model in each architecture direction. Compared with the problem of poor practicability of the trained image recognition model in the prior art, the embodiment can obtain the recognition result of continuous frames by obtaining sample images with continuous frames and taking the sample images as training samples, thereby avoiding the inaccurate recognition caused by recognizing one frame of image. In addition, the embodiment can change the architecture direction of the sample objects in each frame of image, obtain sample direction images in different architecture directions, take the sample direction images in different architecture directions as training samples, obtain the trained image recognition model in different architecture directions, and ensure the accuracy of the recognition result when the architecture direction of the object in the image is different, thereby improving the practicability of the trained image recognition model.
[0172] Optionally, the label acquisition module 1403 comprises:
[0173] The first skeleton determination unit is configured to determine, for each frame of image, the skeleton of the sample object in the image.
[0174] The disease determination unit is configured to determine the skeleton line segment with the disease in the skeleton.
[0175] The label determination unit is configured to determine, in the image, the image region containing the skeleton line segment as the training label.
[0176] Optionally, the model training module 1404 comprises:
[0177] The feature enhancement unit is configured to perform feature enhancement on the corresponding image feature based on the training label in the process of processing the image feature in the sample direction image through the image recognition model, and train the image recognition model according to the enhanced image feature.
[0178] Optionally, the device 1400 further comprises:
[0179] The test image acquisition module is configured to obtain a test image for a test object, the test image including multiple frames of images.
[0180] The test direction image acquisition module is configured to acquire test direction images under different architecture directions by changing the architecture direction of the test object in the current frame test image in the test video.
[0181] The test direction image recognition module is configured to recognize the test direction images under different architecture directions by using the trained image recognition model, and obtain corresponding recognition results.
[0182] The similarity comparison module is configured to compare the test labels of the test direction images under different architecture directions with the corresponding recognition results in terms of similarity, and obtain corresponding similarities of the test direction images under different architecture directions.
[0183] The similarity integration module is configured to integrate the similarities under different architecture directions, and obtain a cumulative similarity corresponding to the current frame test image.
[0184] The model precision determination module is configured to determine that the trained image recognition model reaches a preset recognition precision in a case where the cumulative similarities corresponding to each of the images with a continuous first preset number of frames in the test video all reach a preset threshold.
[0185] Optionally, the similarity comparison module includes:
[0186] The second skeleton determination unit is configured to determine a skeleton of the test object in the test direction image for each test direction image under different architecture directions.
[0187] The dimension determination unit is configured to determine a plurality of dimensions for constructing a vector based on a number of angles formed by line segments composed of the joints in the skeleton.
[0188] The first vector obtaining unit is configured to determine a joint matched with the test label corresponding to the test direction image, determine a value of each dimension in the plurality of dimensions, and enhance the value of the corresponding dimension of the matched joint, to obtain a first vector.
[0189] The second vector obtaining unit is configured to determine a joint matched with the recognition result corresponding to the test direction image, determine a value of each dimension in the plurality of dimensions, and enhance the value of the corresponding dimension of the matched joint, to obtain a second vector.
[0190] The similarity obtaining unit is configured to calculate a similarity between the first vector and the second vector as the similarity corresponding to the test direction image.
[0191] Optionally, the apparatus 1400 further includes:
[0192] The target video acquisition module is configured to acquire real-time target videos for the target object during movement of the target object along the guide route.
[0193] an object image recognition module, configured to, for a current frame object image in a real-time target video, perform recognition on the current frame object image by an image recognition model reaching a preset recognition accuracy, and obtain a recognition result;
[0194] a start interception frame determination module, configured to, in a case where the recognition result is that there is a symptom area, determine the current frame object image as a start interception frame;
[0195] a candidate image frame determination module, configured to, for a subsequent image frame of the start interception frame, determine a candidate image frame in which the subsequent image frame has a recognition result that there is no symptom area;
[0196] an end interception frame determination module, configured to, in a case where, from the candidate image frame, there are a plurality of image frames with recognition results that there is no symptom area, determine a corresponding candidate image frame as an end interception frame;
[0197] a target video interception module, configured to, based on the start interception frame and the end interception frame, intercept the real-time target video, and obtain a target video segment.
[0198] The various modules in the training device of the image recognition model can be realized by software, hardware, or a combination thereof. The various modules can be embedded in or independent of a processor in a computer device in hardware form, or stored in a memory in the computer device in software form, so as to be called and executed by a processor to perform operations corresponding to the various modules.
[0199] In an embodiment, a computer device is provided, which can be a terminal. An internal structure diagram of the computer device can be as shown in Figure 15The computer device shown in the figure includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capability. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be realized through WIFI, mobile cellular network, NFC (near field communication) or other technologies. The computer program is executed by the processor to realize a training method of an image recognition model. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device, or an external keyboard, touchpad or mouse, etc.
[0200] Those skilled in the art can understand that, Figure 15 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0201] In one embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to realize the steps in each of the above method embodiments.
[0202] In one embodiment, a computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to realize the steps in each of the above method embodiments.
[0203] In one embodiment, a computer program product is provided, including a computer program, and the computer program is executed by a processor to realize the steps in each of the above method embodiments.
[0204] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc. The database involved in the embodiments provided in the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in the embodiments provided in the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., without being limited thereto.
[0205] Any combination of the technical features of the above embodiments can be made. In order to make the description simple, all possible combinations of the technical features in the above embodiments are not described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.
[0206] The above embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent of the present application. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A training method for an image recognition model, characterized in that, The method comprises: obtaining a sample image of a sample object, the sample image comprising multiple frames of images; obtaining sample direction images under different skeleton directions by changing the skeleton direction of the sample object in each frame of image; obtaining training labels corresponding to the disease position of the sample object in each frame of image of each sample direction image; training an image recognition model through each sample direction image and the corresponding training label, and obtaining a trained image recognition model under each skeleton direction; The method further comprises: obtaining a test image of a test object, the test image comprising multiple frames of images; for a current frame of test image in the test image, obtaining test direction images under different skeleton directions by changing the skeleton direction of the test object in the current frame of test image; identifying the test direction images under different skeleton directions through the trained image recognition model to obtain the corresponding identification result; comparing the test labels of the test direction images under different skeleton directions with the corresponding identification result to obtain the corresponding similarity of the test direction images under different skeleton directions; integrating the similarities under different skeleton directions to obtain the corresponding cumulative similarity of the current frame of test image; in the case that the corresponding cumulative similarity of each of the images with a continuous first preset number of frames in the test image reaches a preset threshold, determining that the trained image recognition model reaches a preset recognition accuracy; The comparison of the test labels of the test direction images under different skeleton directions with the corresponding identification result to obtain the corresponding similarity of the test direction images under different skeleton directions comprises: for each test direction image under each skeleton direction, determining the skeleton of the test object in the test direction image; determining a plurality of dimensions for constructing a vector based on the number of angles formed by the line segments composed of the joint nodes in the skeleton; determining the joint nodes matched with the corresponding test label of the test direction image, determining the value of each dimension in the plurality of dimensions, and enhancing the value of the corresponding dimension of the matched joint nodes to obtain a first vector; determining the joint nodes matched with the corresponding identification result of the test direction image, determining the value of each dimension in the plurality of dimensions, and enhancing the value of the corresponding dimension of the matched joint nodes to obtain a second vector; calculating the similarity between the first vector and the second vector as the corresponding similarity of the test direction image.
2. The method of claim 1, wherein, The obtaining of the training labels corresponding to the disease position of the sample object in each frame of image of each sample direction image comprises: for each frame of image, determining the skeleton of the sample object in the image; determining the skeleton line segment of the image with the disease in the skeleton; determining the image region containing the skeleton line segment in the image as the training label.
3. The method of claim 1, wherein, The training of the image recognition model through each sample direction image and the corresponding training label comprises: In a process of processing image features in the sample directional image by the image recognition model, the image features are enhanced based on the training labels, and the image recognition model is trained according to the enhanced image features.
4. The method of claim 1, wherein, The method further includes: acquiring a real-time target image for the target object during movement of the target object along the guidance route; identifying a current frame target image in the real-time target image by the image recognition model reaching the preset recognition accuracy, to obtain an identification result; in a case where the identification result is that a symptom area exists, taking the current frame target image as a starting frame of extraction; for a subsequent image frame of the starting frame of extraction, determining a candidate image frame in which the identification result of the subsequent image frame is that a symptom area does not exist; in a case where the identification result of each of image frames existing for a continuous second preset number of frames from the candidate image frame is that a symptom area does not exist, determining the corresponding candidate image frame as an ending frame of extraction; based on the starting frame of extraction and the ending frame of extraction, extracting the real-time target image to obtain a target image segment. 5.A device for training an image recognition model, comprising: The device includes: a sample image acquisition module configured to acquire a sample image for a sample object, the sample image including multiple frames of images; an architecture direction changing module configured to acquire sample directional images in different architecture directions by changing an architecture direction of the sample object in each frame of image; a label acquisition module configured to acquire a training label corresponding to a symptom position of the sample object in each frame of image of each sample directional image; a model training module configured to train an image recognition model by each sample directional image and the corresponding training label, to obtain a trained image recognition model in each architecture direction; The device further includes: a test image acquisition module configured to acquire a test image for a test object, the test image including multiple frames of images; a test directional image acquisition module configured to acquire test directional images in different architecture directions by changing an architecture direction of the test object in a current frame of test image in the test image; a test directional image identification module configured to identify the test directional images in different architecture directions by the trained image recognition model, to obtain corresponding identification results; a similarity comparison module configured to compare test labels of the test directional images in different architecture directions with the corresponding identification results according to similarity, to obtain corresponding similarities of the test directional images in different architecture directions; a similarity integration module configured to integrate the similarities in different architecture directions, to obtain a cumulative similarity corresponding to the current frame of test image; a model accuracy determination module configured to determine that the trained image recognition model reaches a preset recognition accuracy in a case where the cumulative similarity corresponding to each of image frames existing for a continuous first preset number of frames in the test image reaches a preset threshold value; The similarity comparison module includes: The second skeleton determination unit is configured to determine a skeleton of the test object in the test direction image for each architecture direction; The dimension determination unit is configured to determine a plurality of dimensions for constructing a vector based on an angle number formed by a line segment composed of a joint in the skeleton; The first vector obtaining unit is configured to determine a joint matched with a test label corresponding to the test direction image, determine a value of each dimension in the plurality of dimensions, and enhance the value of the corresponding dimension of the matched joint to obtain a first vector; The second vector obtaining unit is configured to determine a joint matched with a recognition result corresponding to the test direction image, determine a value of each dimension in the plurality of dimensions, and enhance the value of the corresponding dimension of the matched joint to obtain a second vector; The similarity obtaining unit is configured to calculate a similarity between the first vector and the second vector as a similarity corresponding to the test direction image.
6. The apparatus of claim 5, wherein, The label obtaining module comprises: The first skeleton determination unit is configured to determine a skeleton of the sample object in the image for each frame of image; The symptom determination unit is configured to determine a skeleton line segment with a symptom in the skeleton; The label determination unit is configured to determine an image region containing the skeleton line segment in the image as a training label.
7. The apparatus of claim 5, wherein, The model training module comprises: The feature enhancement unit is configured to perform feature enhancement on a corresponding image feature based on the training label in a processing process of the image feature in the sample direction image by the image recognition model, and train the image recognition model according to the enhanced image feature.
8. A computer device comprising a memory and a processor, the memory storing a computer program, characterized in that, The processor executes the computer program to realize the steps of the method in any one of claims 1 to 4.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method in any one of claims 1 to 4.
10. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the method in any one of claims 1 to 4.
Citation Information
Patent Citations
Image classification method and device, computer equipment and readable storage medium
CN111160367A