Model training method and device, equipment and readable storage medium
By introducing rotation angle features into the target object detection model and using candidate boxes with rotation angles for marking and defining, the problem of overlapping detection boxes in traditional methods is solved, thus improving detection accuracy.
Patent Information
- Application Number
- CN202210895990.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-26
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Traditional horizontal rectangular bounding box marking methods cause overlapping of detection boxes when detecting target objects with different angles and large aspect ratios, thus reducing detection accuracy.
Candidate boxes with rotation angles are used for labeling and defining. Multiple candidate box information is generated through the candidate box output network and the regression output network. The object detection model is trained by combining multi-dimensional feature information to learn the position and rotation angle features of the target object.
It improves the accuracy of target object detection, reduces the overlap rate between multiple target object detection boxes, and enhances the detection effect of the model.
Smart Images

Figure CN115205634B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to a model training method, a processing device, a computer device and a computer readable storage medium. BACKGROUND
[0002] With the continuous development of image processing technology, more and more scenes need to detect target objects in the to-be-detected image and mark the detected target objects. At present, when target detection is performed, a standard horizontal marking box is used to frame a target object to complete the marking of the target object. For target objects with different angles and large length-width ratios, a large amount of overlap will occur between the marking boxes, thereby reducing the accuracy of target object detection. SUMMARY
[0003] The present application provides a model training method, device, equipment and readable storage medium, which is beneficial to output the position and rotation angle of the target object in the image by the target object detection model, and improves the accuracy of target object detection.
[0004] In a first aspect, the present application provides a model training method, which comprises:
[0005] obtaining a sample image, and obtaining first marking box information corresponding to the sample image, wherein the first marking box information is used to indicate the position of a first target object in the sample image and the rotation angle of the first target object relative to a reference axis;
[0006] inputting the sample image into the object detection model to obtain a plurality of candidate box information corresponding to the sample image output by the candidate box output network and a plurality of prediction box information corresponding to the plurality of candidate box information output by the regression output network, wherein the candidate box information is used to indicate the candidate position and candidate rotation angle of the target object output by the object detection model, and the prediction box information is obtained by performing position regression on the candidate box information;
[0007] determining first prediction box information in the plurality of prediction box information according to the plurality of candidate box information and the first marking box information, wherein the coincidence degree between the candidate box corresponding to the first candidate box information corresponding to the first prediction box information and the marking box corresponding to the first marking box information is greater than a preset threshold;
[0008] training the object detection model according to a first relationship parameter between the first prediction box information and the first candidate box information and a second relationship parameter between the first marking box information and the first candidate box information to obtain a target object detection model.
[0009] In a second aspect, the present application provides an image processing method, which comprises:
[0010] obtaining a to-be-detected image;
[0011] performing object detection processing on the to-be-detected image by a target object detection model to determine second bounding box information of a second target object in the to-be-detected image, the second bounding box information being used to indicate a position of the second target object in the to-be-detected image and a rotation angle of the second target object relative to a reference axis, the target object detection model being obtained by the model training method of the first aspect;
[0012] adding a bounding box for the second target object in the to-be-detected image according to the second bounding box information.
[0013] In a third aspect, the present application provides a processing apparatus, which comprises a module for implementing the model training method or a module for implementing the image processing method.
[0014] In a fourth aspect, the present application provides a computer device, which comprises a processor, a storage device and a communication interface, wherein the processor, the communication interface and the storage device are connected to each other, the storage device stores executable program codes, the processor is configured to invoke the executable program codes to implement the module for implementing the model training method or the module for implementing the image processing method.
[0015] In a fifth aspect, the present application provides a computer readable storage medium, which stores a computer program, the computer program comprises program instructions, and the program instructions are executed by a processor to implement the model training method or the image processing method.
[0016] In a sixth aspect, the present application provides a computer program product, which comprises a computer program or computer instructions, and the computer program or computer instructions are executed by a processor to implement the model training method or the image processing method.
[0017] The method provided in the application first acquires a sample image and first mark box information indicating a position of a first target object in the sample image and a rotation angle of the first target object relative to a reference axis; then inputs the sample image into an object detection model to obtain a plurality of candidate box information corresponding to the sample image output by a candidate box output network and a plurality of prediction box information corresponding to the plurality of candidate box information output by a regression output network; then determines first prediction box information in the plurality of prediction box information based on a coincidence degree, and trains the object detection model according to a first relationship parameter between the first prediction box information and first candidate box information and a second relationship parameter between the first mark box information and the first candidate box information, so as to combine multi-dimensional feature information and ensure the training effect of the model, and finally obtain a target object detection model. The application introduces the rotation angle as a feature in the training process, so that the model can learn the rotation angle feature of the first target object in the sample image. Compared with the traditional horizontal rectangular box detection method, the target object detection model obtained by the model training method provided in the application is beneficial to the target object detection model to output the position and rotation angle of the target object in the image, so that the detection result can be more matched with the real position and rotation angle of the target object in the image, thereby improving the accuracy of target object detection. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed in the embodiment description will be briefly introduced as follows. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.
[0019] Figure 1 is an effect diagram of target object detection provided by an example embodiment of the application;
[0020] Figure 2 is an architecture diagram of an image processing system provided by an example embodiment of the application;
[0021] Figure 3 is a flow diagram of a model training method provided by an example embodiment of the application;
[0022] Figure 4A is an effect diagram of target object detection provided by another example embodiment of the application;
[0023] Figure 4B is a network structure diagram of an object detection model provided by an example embodiment of the application;
[0024] Figure 4C is a diagram of generating a candidate box provided by an example embodiment of the application;
[0025] Figure 4D is an effect diagram of target object detection by a target detection model provided by an example embodiment of the present application;
[0026] Figure 5 is a flow diagram of an image processing method provided by an example embodiment of the present application;
[0027] Figure 6 is a schematic block diagram of a processing device provided by an example embodiment of the present application;
[0028] Figure 7 is a schematic block diagram of a computer device provided by an example embodiment of the present application. DETAILED DESCRIPTION
[0029] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0030] It should be noted that the "first", "second", and the like described in the embodiments of the present application are only for the purpose of description, and cannot be understood as indicating or implying the relative importance of the technical features or implicitly indicating the number of the indicated technical features. Therefore, the technical features limited by "first" and "second" can explicitly or implicitly include at least one of the features.
[0031] With the development of information systems, users can collect a to-be-detected image by an image acquisition device, and then detect target objects in the to-be-detected image by a target detection method, and mark the detected target objects. In a traditional target detection algorithm, a standard horizontal rectangular frame is used for bounding the target objects, but for target objects with different angles and large length-width ratios (such as long strip-shaped pigs), this bounding method will cause a large amount of overlap between the detection frames, thereby reducing the detection accuracy. An effect diagram of target object detection by a traditional bounding method is shown in Figure 1 . The diagram includes multiple pigs, and by a traditional target detection algorithm, a horizontal rectangular frame can be marked for the position of each pig in the to-be-detected image. Since the target objects are unevenly distributed in the to-be-detected image (may be scattered or clustered), when the target objects are clustered, the detection frames corresponding to multiple target objects will overlap. For example, in Figure 1In the figure, three pigs distributed in an inclined manner are framed by a horizontal rectangular frame, and a large amount of overlap occurs between the detection frames of the three pigs. In the figure, the black points are boundary points of the marked pigs, and the horizontal rectangular frame is the detection frame of the marked pigs. Training the model based on the training data marked by the horizontal rectangular frame affects the training effect of the model. Moreover, using the model trained by the training data to detect the target object in the to-be-detected image causes the detected target object in the to-be-detected image to be marked in the form of a horizontal rectangular frame, which reduces the detection accuracy.
[0032] Based on the defects of the above method, the present application proposes a target object detection model to realize the detection of a target object based on the characteristics of the target object in a top view, different shapes, and a large length-width ratio of the body. The target object is framed by a rotatable marking frame. Moreover, in the training stage, the candidate frame is set as a candidate frame with a rotation angle, which enables the target object detection model to learn the position and rotation angle of the marking frame and other features, and is beneficial to the target object detection model to output the position and rotation angle of the target object in the image, thereby improving the accuracy of target object detection.
[0033] It can be understood that in the specific embodiments of the present application, related data such as to-be-detected images and sample images are involved. When the above embodiments of the present application are applied to specific products or technologies, the collection, use, and processing of related data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0034] The present application will be specifically described through the following embodiments:
[0035] Please refer to Figure 2 The figure is an architecture schematic diagram of an image processing system provided by an example embodiment of the present application. The image processing system specifically can include a terminal device 201 and a server 202. The terminal device 201 and the server 202 are connected through a network, such as a local area network, a wide area network, a mobile Internet, etc. An operation object operates on a browser or a client application of the terminal device 201 to perform detection operations on various image data. The server 202 can respond to the operation to provide various image data detection services for the operation object.
[0036] In an embodiment, the terminal device 201 can be an image acquisition device (for example, a camera installed on the top of a pig house), and the server 202 includes a target object detection model. The terminal device 201 can acquire (for example, real-time acquisition, interval acquisition) a to-be-detected image including a target object; the server 202 obtains the to-be-detected image from the terminal device 201, and performs detection processing on the target object in the to-be-detected image to obtain a detection result of the target object (for example, adding a mark box to the target object in the to-be-detected image); and the server 202 can perform processing and analysis operations on the target object according to the detection result of the target object (for example, target object quantity statistics, target object distribution analysis, etc.).
[0037] In an embodiment, the terminal device 201 can be a computer device, and the terminal device 201 stores a to-be-detected image including a target object, and the server 202 includes a target object detection model. The server 202 obtains the to-be-detected image from the terminal device 201, and performs detection processing on the target object in the to-be-detected image to obtain a detection result of the target object (for example, adding a mark box to the target object in the to-be-detected image); and the server 202 returns the detection result of the target object to the terminal device 201, so that the terminal device 201 performs processing and analysis operations on the target object based on the returned result (for example, target object quantity statistics, target object distribution analysis, etc.).
[0038] The terminal device 201 is also referred to as a terminal, a user equipment (UE), an access terminal, a user unit, a mobile device, a user terminal, a wireless communication device, a user agent, or a user apparatus. The terminal device can be a smart home appliance, a handheld device (for example, a smartphone, a tablet computer) with a wireless communication function, a computing device (for example, a personal computer (PC), a vehicle-mounted terminal, a smart voice interaction device, a wearable device, or other smart devices, etc.), but is not limited thereto.
[0039] The server 202 can be a standalone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, etc. basic cloud computing services.
[0040] In an embodiment, the architecture of the image processing system provided in the present application can further include a database for storing the to-be-detected image, the detection result of the target object, the label box information, and the like, and can also be used to store the related data of the target object detection model. These data can be recorded in different database tables in the database. For example, the database can be a database provided in a server, i.e., a built-in or self-provided database of the server. The database can also be an external database connected with the server, such as a cloud database (i.e., a database deployed in the cloud). The cloud database can be deployed based on any one of a private cloud, a public cloud, a hybrid cloud, an edge cloud, and the like, so that the cloud database focuses on different functions. For example, the database deployed in the private cloud is based on the user's own device, and focuses on serving a small number of users. The database deployed in the public cloud is based on a third-party provided cloud platform, and can enable the data stored in the database to be shared, so that the data of any user can be stored in the database, and the data in the database can be used by any user.
[0041] It can be understood that the architecture schematic diagram of the system described in the embodiments of the present application is used to more clearly illustrate the technical solutions of the embodiments of the present application, and does not constitute a limitation on the technical solutions provided by the embodiments of the present application. For example, the method provided by the embodiments of the present application can be executed by the server 202, and can also be executed by other servers or server clusters different from the server 202 and capable of communicating with the terminal device 201 and / or the server 202. Those skilled in the art can know that the number of terminal devices and servers in Figure 2 the above is only schematic. According to the needs of business implementation, any number of terminal devices and servers can be configured. Moreover, as the system architecture evolves and new business scenarios appear, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems. In subsequent embodiments, the terminal device will be referred to as the terminal device 201, and the server will be referred to as the server 202, which will not be described in detail in subsequent embodiments.
[0042] Please refer to Figure 3 , Figure 3 is a flowchart of a model training method provided by an exemplary embodiment of the present application. The method is used to illustrate a server for detecting a target object (for example, the server 202 in Figure 2 ), which can include the following steps:
[0043] S301, a sample image is obtained, and first label box information corresponding to the sample image is obtained. The first label box information is used to indicate the position of a first target object in the sample image and the rotation angle relative to a reference axis.
[0044] In the embodiments of the present application, the sample image is an image containing at least one first target object. For each first target object contained in the sample image, a first marking box information corresponds, and the first marking box information is used to indicate the position of the first target object in the sample image and the rotation angle relative to the reference axis. That is, the first marking box information can represent the rotation angle formed by the posture of the first target object in the sample image and the reference axis (for example, the horizontal reference axis or the vertical reference axis based on the sample image frame). Based on the above first marking box information including the rotation angle feature as the training data of the target object detection model, the model training is carried out, so that the model can learn the rotation angle feature of the target object in the sample image, which is beneficial to the target object detection model to output the rotation angle of the target object in the image. Compared with the traditional horizontal rectangular frame detection method, the present application can reduce the overlap rate between the detection frames corresponding to multiple target objects, so as to improve the accuracy of target object detection.
[0045] In the present application, each marking box information corresponds to a marking box in the sample image. Since the first marking box information includes the parameter for indicating the rotation angle of the first target object relative to the reference axis, that is, the marking box corresponding to each marking box information can be displayed at a certain rotation angle in the sample image (for example, by a rectangular frame that can have an arbitrary rotation angle with the reference axis).
[0046] In an embodiment, the first target object can be a person, an animal (such as a pig), a plant, etc. in the image. The sample image can be a top view image of the first target object, or an image of the first target object at an arbitrary angle with the reference plane.
[0047] In an embodiment, the above process of obtaining the first marking box information corresponding to the sample image can include the following steps (a1-a2):
[0048] (a1), marking the boundary points of the first target object in the sample image, and calculating the minimum circumscribed rectangle of the boundary points of the first target object.
[0049] (a2), determining the position parameter of the minimum circumscribed rectangle as the first marking box information corresponding to the sample image, and the position parameter of the minimum circumscribed rectangle includes the position of the minimum circumscribed rectangle in the sample image and the rotation angle of the minimum circumscribed rectangle relative to the reference axis.
[0050] In an embodiment, the present application can be applied to the scenario of pig detection. In this scenario, the first target object can be a pig, and the sample image can be an image containing a pig. The sample image can be obtained by image acquisition of a pigsty through a camera installed on the top of the pigsty. The first mark box information is obtained by marking the sample image. The steps (a1-a2) will be described below based on the above scenario.
[0051] In step (a1-a2), the marking method for marking the sample image is as follows: first, boundary points of each pig in the sample image are marked, and the number of boundary points is multiple, which can indicate the approximate outline and posture of the pig in the sample image to a large extent; second, for each pig in the sample image, the minimum circumscribed rectangle of the boundary points of each pig is calculated, and the minimum circumscribed rectangle is regarded as the first mark box corresponding to the pig; third, for the minimum circumscribed rectangle corresponding to each pig in the sample image, the center point coordinates (x, y) of the minimum circumscribed rectangle, the width (w) and height (h) of the minimum circumscribed rectangle, and the counterclockwise angle θ (0 < θ < π) of the minimum circumscribed rectangle relative to the horizontal direction (i.e. reference axis) (i.e. rotation angle) are determined, so as to obtain the first mark box information (x, y, w, h, θ) corresponding to the first mark box. The rotation angle mentioned above is the rotation angle of the minimum circumscribed rectangle relative to the reference axis. As shown in the figure, the figure is an effect diagram of target object detection based on the method provided by the present application, which includes a plurality of pigs, and the black points in the figure are the boundary points of the marked pigs, and the rectangular box with a rotation angle is the first mark box of the marked pigs. Compared with the horizontal rectangular box obtained by the traditional target detection algorithm, the rotation angle of the first mark box relative to the reference axis can be arbitrary, which reduces the overlap rate between the detection boxes corresponding to multiple first target objects. Figure 4A
[0052] In order to ensure the uniqueness of the first mark box, when marking the first mark box, if the rotation angle is greater than π / 2, the rotation angle is marked as θ-π / 2, and w and h are replaced, i.e. the first mark box is expressed by the first mark box information (x, y, h, w, θ-π / 2). Through the above method, the calculation error caused by inaccurate marking of the rotation angle of the first mark box is avoided, and the training effect of the object detection model is ensured.
[0053] S302, input the sample image into the object detection model to obtain a plurality of candidate box information corresponding to the sample image output by the candidate box output network and a plurality of prediction box information corresponding to the plurality of candidate box information output by the regression output network, the candidate box information is used to indicate the candidate position and candidate rotation angle of the first target object output by the object detection model, and the prediction box information is obtained by position regression on the candidate box information.
[0054] In an embodiment of the present application, the object detection model comprises a candidate box output network and a regression output network. The candidate box output network is configured to generate a plurality of candidate box information in the sample image, each candidate box information corresponding to a candidate box. The regression output network is configured to perform position regression processing on the candidate box information to obtain a plurality of prediction box information corresponding to the plurality of candidate box information (i.e., to obtain a plurality of prediction boxes corresponding to the plurality of candidate boxes). Each candidate box uniquely corresponds to a prediction box. The difference between the prediction box and the candidate box is that the candidate box information of the candidate box is pre-set, so that the candidate box can only present a fixed plurality of shapes (for example, each candidate box can only be any combination of a plurality of pre-set sizes, a plurality of aspect ratios, and a plurality of rotation angles). The prediction box is obtained by performing position regression processing on the candidate box information. The position regression processing can be regarded as fine-tuning the candidate box, so that the prediction box can more accurately and more finely present the shape of the first target object (for example, each prediction box can be any combination of any size, any aspect ratio, and any rotation angle). By performing position regression on the candidate box to obtain the prediction box, the prediction box can be more matched with the real position and rotation angle of the first target object in the image, rather than being limited to a fixed plurality of sizes, a plurality of aspect ratios, and a plurality of rotation angles. Subsequent model training based on the prediction box obtained by performing position regression on the candidate box can improve the accuracy of model prediction.
[0055] In an embodiment, the object detection model of the present application further comprises a feature extraction network. The feature extraction network is configured to perform feature extraction on the sample image to obtain a feature map corresponding to the sample image.
[0056] Based on this embodiment, the step of outputting the plurality of candidate box information corresponding to the sample image by the candidate box output network will be described as follows: first, performing feature extraction on the sample image by the feature extraction network to obtain a plurality of feature maps corresponding to the sample image; second, for any one of the plurality of feature maps, mapping each feature point in the feature map to a pixel point in the sample image (i.e., determining the pixel point corresponding to the target feature point in the sample image from the sample image, the target feature point being any one of the feature points included in any one of the plurality of feature maps); and third, for the pixel point to which any one of the feature points is mapped, outputting a plurality of candidate box information corresponding to the pixel point by the candidate box output network, each candidate box information corresponding to a candidate box. Through the above first to third steps, the plurality of candidate box information is obtained.
[0057] For any one feature point mapped to a pixel point of the sample image, the candidate box output network is used to obtain multiple candidate box information corresponding to the pixel point, and the size, aspect ratio and rotation angle corresponding to each candidate box information are not completely the same. Specifically, the feature extraction network is used to perform feature extraction processing on the sample image to obtain multiple feature maps corresponding to the sample image, and the size of each feature map can be different. For any one feature point in the feature map, multiple candidate boxes with different sizes, aspect ratios and rotation angles can be generated based on the size of the feature map at the pixel point corresponding to the sample image.
[0058] In an embodiment, the detection accuracy and detection speed are comprehensively considered, and a RetinaNet network is used as a main network of an object detection model, and the network structure is as shown in Figure 4B The feature extraction network (for example, a Resnet50 network) is used to perform feature extraction processing on the sample image to obtain feature maps at different stages of the sample image (including five feature maps C1, C2, C3, C4 and C5 in the figure). The above step can be regarded as a forward process of the convolutional network, and the size of the multiple feature maps generated in the forward process will change (for example, the sizes of the feature maps C1 and C2 are the same, and the sizes of the feature maps C2, C3, C4 and C5 are different). Then, the feature pyramid structure is established by using the feature maps at different stages of the sample image (three feature maps C3, C4 and C5 are selected in this embodiment), that is, the 2d convolution kernel included in the feature pyramid network is used to perform convolution processing and up-sampling processing on the three feature maps C3, C4 and C5 to obtain multiple feature maps (for example, five feature maps P3, P4, P5, P6 and P7 in the figure) processed by the feature pyramid network. Among them, the feature maps P3, P4 and P5 correspond to C3, C4 and C5 generated by the forward process of the convolutional network, respectively. Since the low-level feature map is not rich in semantics and cannot be directly used for prediction and classification, the deep feature is more reliable. By using the above method of constructing the feature pyramid network, feature maps with different resolutions can be obtained, and they all contain semantic information of the original deepest feature map. By processing the feature maps with different resolutions, the low-resolution feature map with strong semantic information and the high-resolution feature map with weak semantic information but rich spatial information can be fused under the premise of increasing less calculation amount, so as to ensure the model training effect.
[0059] In the object detection model, the candidate box output network is configured to generate a plurality of candidate boxes in the sample image based on the mapping points, which are pixel points in the sample image mapped from each feature point in the plurality of feature maps processed by the feature pyramid network. In other words, the candidate box output network can generate a corresponding candidate box in the sample image for any feature point in the P3, P4, P5, P6, and P7 feature maps.
[0060] The classification regression network in the regression output network can process each candidate box to output a confidence score of each candidate box. Meanwhile, each candidate box corresponds to a prediction box, and the confidence score of each candidate box is consistent with the confidence score of the prediction box corresponding to the candidate box. The position regression network in the regression output network is configured to perform position regression processing on the candidate box information to obtain a plurality of prediction box information corresponding to the plurality of candidate box information (i.e., a plurality of prediction boxes corresponding to the plurality of candidate boxes). Through the feature extraction network and the feature pyramid network, multi-level feature fusion and target detection processing can be achieved, the multi-dimensional feature information of the image can be fully utilized, and the accuracy of target detection is improved.
[0061] For example, for a certain feature point in any one of the plurality of feature maps (e.g., P3, P4, P5, P6, and P7, which are arranged in ascending order of feature map size), a size of a candidate box (e.g., 8*8*2 0 , 8*8*2 1 / 3 , and 8*8*2 2 / 3 ) can be set first; then a width-height ratio (e.g., 1:2, 1:4, 4:1, and 2:1) can be set, and a rotation angle (e.g., 0 degrees, 30 degrees, and 60 degrees) can be set. Based on this, a*b*c candidate boxes can be generated for a certain feature point in any one of the plurality of feature maps, i.e., 3*4*3=36 candidate boxes are generated. Each feature map will generate 36*w i *h i candidate boxes (i=3, 4, 5, 6, and 7), where w i , h i represent the width and height of the i-th feature map, and w i *h i is the number of feature points included in the feature map.
[0062] The above-mentioned setting of the size of the candidate boxes can be achieved through the following steps: For the five feature maps P3, P4, P5, P6, and P7, the basic size of the candidate boxes corresponding to multiple feature maps can be set first (for example, the basic sizes of the five feature maps P3, P4, P5, P6, and P7 are [8*8, 16*16, 32*32, 64*64, 128*128], respectively). Each feature map includes different feature information. For example, feature map P3 is a low-resolution feature map with strong semantic information, and feature map P7 is a high-resolution feature map with weak semantic information but rich spatial information; then, multiple scales (for example, [2 0 ,2 1 / 3 ,2 2 / 3 Based on the basic dimensions of the candidate boxes corresponding to multiple feature maps and multiple scales, the dimensions of multiple candidate boxes can be determined (for example, for any feature point in feature map P3, the size of the candidate box to be generated can be 8*8*2). 0 8*8*2 1 / 3 8*8*2 2 / 3 The dimensions of the three candidate bounding boxes; for any feature point in feature map P4, the size of the candidate bounding box to be generated can be 16*16*2. 0 16*16*2 1 / 3 16*16*2 2 / 3 (Dimensions of the three candidate boxes).
[0063] like Figure 4C As shown in the figure, this diagram illustrates the generation of candidate bounding boxes at a specific feature point in a feature map. The figure includes boxes 401 and 402. Box 401 shows multiple horizontal rectangles generated based on a traditional object detection algorithm. The rotation angle between these horizontal rectangles and the reference axis is a fixed value, making it impossible to adjust the rotation angle based on actual conditions. Box 402 shows candidate bounding boxes with rotation angles generated based on the method proposed in this application. By pre-setting candidate bounding box information (including size parameters, aspect ratio parameters, and rotation angle parameters), the candidate bounding boxes can be generated in various sizes (e.g., 8*8*2). 0 8*8*2 1 / 3 8*8*2 2 / 3 Presented in the form of three candidate box sizes), multiple aspect ratios (e.g., 1:2, 1:4, 4:1, 2:1), and multiple rotation angles (e.g., 0 degrees, 30 degrees, 60 degrees).
[0064] In an embodiment, the regression output network is used for position regression processing of the candidate box information to obtain a plurality of prediction box information corresponding to the plurality of candidate box information. The plurality of prediction box information includes a plurality of groups of prediction box information, and each group of prediction box information includes a plurality of prediction box information corresponding to one feature point in one feature map. Any two prediction box information in the plurality of prediction box information corresponding to one feature point are unmatched, which means that any two prediction box information corresponding to two prediction boxes have one or more different parameters, such as size parameter, aspect ratio parameter and rotation angle parameter. That is, any two candidate box information in the plurality of candidate box information corresponding to one feature point are unmatched, and each candidate box is processed by position regression to obtain a corresponding prediction box, so that any two prediction box information in the plurality of prediction box information corresponding to one feature point are unmatched. Wherein, the regression output network and the classification regression network are separately performed between each feature map, and the network weight is shared to output the candidate box type of the candidate box included in each feature map and the prediction box corresponding to the candidate box.
[0065] In an embodiment, the method for obtaining the prediction box by using the regression output network to perform position regression processing on the candidate box information can be processed by using a position regression formula. The position regression formula is as follows:
[0066] x predicet =p x *width a +x a
[0067] y predicet =p y *height a +y a
[0068]
[0069]
[0070] θ predicet =arctan(tanθ a +p θ )
[0071] Wherein, x a , y a , width a , height a , tanθ a are the center point horizontal coordinate, the center point vertical coordinate, the width, the height and the rotation angle of the candidate box in the candidate box information, respectively. x predicet , y predicet , width predicet , heightpredicet , θ predicet , θ x , θ y , θ w , θ h , θ θ are regression parameters which need to be trained by the regression output network. The training method of the regression parameters will be described in detail in subsequent embodiments, which will not be described here. Through the above method, the position regression formula can be used to determine the prediction box information corresponding to the candidate box information.
[0072] S303, according to the plurality of candidate box information and the first mark box information, determine the first prediction box information in the plurality of prediction box information, the coincidence degree of the candidate box corresponding to the first candidate box information corresponding to the first prediction box information and the mark box corresponding to the first mark box information is greater than a preset threshold.
[0073] In the embodiment of the application, the coincidence degree of the candidate box corresponding to the plurality of prediction box information and the mark box can be used to determine the first prediction box information in the plurality of prediction box information. The number of the first prediction box corresponding to the first prediction box information can be one or more, and one or more first prediction boxes can indicate the position and rotation angle of the first target object in the sample image to a certain extent. Since the first prediction box information is determined from the plurality of prediction box information, the difference between the first prediction box and the prediction box is that the prediction box is one-to-one corresponding to the candidate box, and for each feature point in the feature map, a plurality of candidate boxes will be generated in the sample image, resulting in that the prediction box corresponding to the candidate box has no relevance with the position of the first target object in the sample image (that is, any prediction box does not necessarily represent the position and rotation angle of the first target object in the sample image). The first prediction box information is determined based on the coincidence degree of the candidate box corresponding to the plurality of prediction box information and the mark box, and the prediction box satisfying the coincidence degree condition in the plurality of prediction box information, that is, the first prediction box corresponding to the first prediction box information is the prediction box partially coinciding with the first target object in the sample image, and any first prediction box can represent the position and rotation angle of the first target object in the sample image to a certain extent. Training the object detection model based on the first prediction box improves the model training effect and improves the prediction accuracy of the model.
[0074] In an embodiment, the above process of determining the first prediction box information in the plurality of prediction box information according to the plurality of candidate box information and the first mark box information can include the following steps (b1-b3):
[0075] (b1) determining, according to the plurality of candidate box information and the first marked box information, an overlap degree between each candidate box corresponding to the plurality of candidate box information and the marked box corresponding to the first marked box information.
[0076] The overlap degree between the candidate box and the marked box can be calculated by using an Intersection over Union (IoU) method. The IoU calculates an overlap rate of two bounding boxes, that is, a ratio of an intersection to a union of the two bounding boxes. In this application, the ratio of the intersection to the union of the candidate box and the marked box is calculated.
[0077] (b2) determining the candidate box with the overlap degree greater than a preset threshold as the first candidate box.
[0078] The preset threshold can be set in advance (for example, the preset threshold is set to 0.5). If the overlap degree between a candidate box in the plurality of candidate boxes and the first marked box is greater than or equal to the preset threshold (for example, IoU≥0.5), the candidate box is determined as the first candidate box; if the overlap degree between a candidate box in the plurality of candidate boxes and the first marked box is less than the preset threshold (for example, IoU<0.5), the candidate box is not the first candidate box. It should be noted that the preset threshold can be flexibly set according to actual business conditions. For example, if it is required that a small part of the overlap area between the candidate box and the marked box is regarded as the first candidate box, the preset threshold can be set to a small value (for example, the preset threshold is set to 0.4); if it is required that a large part of the overlap area between the candidate box and the marked box is regarded as the first candidate box, the preset threshold can be set to a large value (for example, the preset threshold is set to 0.8).
[0079] (b3) determining the prediction box information corresponding to the prediction box corresponding to the first candidate box as the first prediction box information.
[0080] The first candidate box is selected from the plurality of candidate boxes, and each candidate box in the plurality of candidate boxes corresponds to a prediction box. Therefore, each first candidate box also corresponds to a prediction box. Then, the prediction box information corresponding to the prediction box corresponding to the first candidate box is determined as the first prediction box information.
[0081] S304, training the object detection model according to a first relationship parameter between the first prediction box information and the first candidate box information, and a second relationship parameter between the first marked box information and the first candidate box information, to obtain a target object detection model.
[0082] In the embodiments of the present application, the object detection model is trained by using the first prediction box information, the first candidate box information and the first marked box information, so as to combine multi-dimensional feature information and ensure the training effect of the model.
[0083] In an embodiment, the regression output network comprises a position regression network and a classification regression network, the position regression network being configured to output a plurality of pieces of predicted bounding box information, and the classification regression network being configured to output a plurality of pieces of confidence corresponding to the plurality of pieces of predicted bounding box information.
[0084] The classification regression network can process each candidate bounding box to output a confidence of the candidate bounding box. Meanwhile, each candidate bounding box corresponds to a predicted bounding box, and the confidence of each candidate bounding box is consistent with the confidence of the predicted bounding box corresponding to the candidate bounding box.
[0085] The process of training the object detection model according to the first relationship parameter between the first predicted bounding box information and the first candidate bounding box information and the second relationship parameter between the first labeled bounding box information and the first candidate bounding box information to obtain the target object detection model can include the following steps (c1-c3):
[0086] (c1) calculating a first loss of the object detection model according to the first relationship parameter between the first predicted bounding box information and the first candidate bounding box information and the second relationship parameter between the first labeled bounding box information and the first candidate bounding box information, the first loss being configured to indicate an accuracy of the object detection model in locating the position of the first target object.
[0087] In the embodiment, the first loss of the object detection model is configured to indicate the accuracy of the object detection model in locating the position of the first target object, that is, a loss corresponding to the position regression network in the object detection model.
[0088] The first relationship parameter can be obtained by the first predicted bounding box information, the first candidate bounding box information and the position regression formula in the foregoing embodiment, and the second relationship parameter can be obtained by the first labeled bounding box information, the first candidate bounding box information and the position regression formula in the foregoing embodiment. It can be understood that, in the foregoing embodiment, the position regression formula p x , p y , p w , p h , p θ With the model training being determined, the position regression formula is used to determine the predicted bounding box corresponding to the candidate bounding box. In the embodiment, how to determine the regression parameter in the position regression formula by using the first predicted bounding box information, the first candidate bounding box information and the first labeled bounding box information is described, that is, a training method of the regression parameter.
[0089] For example, the calculation formula of the first loss is as follows:
[0090]
[0091] x = y1-y2
[0092] y1 = (p 1x , p1y ,p 1w ,p 1h ,p 1θ )
[0093] y2 = (p 2x ,p 2y ,p 2w ,p 2h ,p 2θ )
[0094] wherein, smooth L is a first loss (i.e. position regression loss), (p 1x ,p 1y ,p 1w ,p 1h ,p 1θ ) is a first relationship parameter between the first predicted bounding box information and the first candidate bounding box information, (p 2x ,p 2y ,p 2w ,p 2h ,p 2θ ) is a second relationship parameter between the first labeled bounding box information and the first candidate bounding box information.
[0095] (p 1x ,p 1y ,p 1w ,p 1h ,p 1θ ) can be determined based on the first predicted bounding box information, the first candidate bounding box information and a position regression formula; (p 2x ,p 2y ,p 2w ,p 2h ,p 2θ ) can be determined based on the first labeled bounding box information, the first candidate bounding box information and the position regression formula.
[0096] For example, for any set of training data (including any one first predicted bounding box information, first candidate bounding box information corresponding to the first predicted bounding box information, first labeled bounding box information), the center point horizontal coordinate, center point vertical coordinate, width, height and rotation angle of the first predicted bounding box in the first predicted bounding box information can be taken as x predicet , y predicet , width predicet , height predicet and θ predicet in the position regression formula; the center point horizontal coordinate, center point vertical coordinate, width, height and rotation angle of the first candidate bounding box in the first candidate bounding box information can be taken as x a , y a , width a , height a and tanθa The first candidate box information and the first prediction box information are substituted into the position regression formula, and a value (for example, y1) of a set of regression parameters corresponding to the set of training data can be obtained.
[0097] The center point horizontal coordinate, the center point vertical coordinate, the width, the height and the rotation angle of the first bounding box in the first bounding box information are taken as x predicet , y predicet , width predicet , height predicet and θ predicet in the position regression formula. The center point horizontal coordinate, the center point vertical coordinate, the width, the height and the rotation angle of the first bounding box in the first bounding box information are taken as x a , y a , width a , height a and tanθ a in the position regression formula. The first candidate box information and the first bounding box information are substituted into the position regression formula, and a value (for example, y2) of a set of regression parameters corresponding to the set of training data can be obtained. At this time, the first relationship parameter and the second relationship parameter have been obtained, and the object detection model can be adjusted. A plurality of sets of first relationship parameters and second relationship parameters can be obtained through a plurality of sets of training data, and the object detection model can be iteratively adjusted through the plurality of sets of first relationship parameters and second relationship parameters until the object detection model converges, and a target object detection model is obtained. The first loss is calculated through the first relationship parameter and the second relationship parameter, so that the first loss can sufficiently learn the difference information between the first prediction box information and the first candidate box information, and the difference information between the first bounding box information and the first candidate box information. In this way, the training effect of the model is ensured. The object detection model is trained based on the first loss, and a target object detection model is obtained, so as to improve the accuracy of the target object detection model in target object detection.
[0098] In an embodiment, the regression parameter of the position regression network at the current time can be taken as the first relationship parameter between the first prediction box information and the first candidate box information. In this embodiment, only the second relationship parameter between the first bounding box information and the first candidate box information is calculated, and the object detection model is adjusted through the calculation formula of the first loss, so as to reduce the calculation amount and improve the training efficiency.
[0099] (c2), according to the confidence of each prediction box information, a second loss of the object detection model is calculated, and the second loss is used to indicate the accuracy of the classification of the object detection model.
[0100] In the embodiments of the present application, the second loss of the object detection model is used to indicate the accuracy of the classification of the object detection model, that is, the loss corresponding to the classification regression network in the object detection model. The confidence of each of the plurality of prediction box information corresponds to the confidence of the candidate box corresponding to the plurality of prediction box information, and the confidence of the candidate box is obtained by processing the candidate box by the classification regression network. In the present application, the reference confidence can be determined according to the overlap between the prediction box corresponding to each prediction box information and the first marking box corresponding to the first marking box information, and then the second loss of the object detection model can be calculated according to the confidence of each of the plurality of prediction box information and the reference confidence calculated by the above method.
[0101] For example, the calculation formula of the second loss is as follows:
[0102] FL(P T )=-α t (1-p t ) γ log(p t )
[0103]
[0104] Wherein, FL(P T ) is the second loss (i.e. classification loss), the value of a is generally 0.25, the value of g is generally 2, p is the confidence of the prediction box, the confidence of the prediction box is equal to the confidence of the candidate box corresponding to the prediction box, and the confidence of the candidate box is obtained by processing the candidate box by the classification regression network. K is the sample type of the prediction box. When k = 1, it means that the sample type of the candidate box corresponding to the prediction box is positive sample type (i.e. the overlap between the candidate box and the marking box is greater than or equal to the preset threshold). When k≠1, it means that the sample type of the candidate box corresponding to the prediction box is negative sample type (i.e. the overlap between the candidate box and the marking box is less than the preset threshold).
[0105] The calculation formula of the second loss can be used to determine a plurality of second losses by using the confidence of each of the plurality of prediction box information and the reference confidence, and the object detection model is iteratively adjusted according to the plurality of second losses, until the object detection model converges, and the target object detection model is obtained.
[0106] (c3), iteratively adjusting the object detection model according to the first loss and the second loss to obtain a target object detection model.
[0107] In an embodiment, the process of step (c3) of iteratively adjusting the object detection model according to the first loss and the second loss to obtain a target object detection model can include the following steps (c31-c32):
[0108] (c31) calculating a total loss of the object detection model according to the first loss, the second loss, and the number of the first bounding box information.
[0109] For example, the calculation formula of the total loss is as follows:
[0110]
[0111] wherein, smooth L is the first loss, FL(P T ) is the second loss, and s is the number of the first bounding box information.
[0112] (c32) iteratively adjusting parameters of the object detection model according to the total loss to obtain a target object detection model.
[0113] In the embodiments of the present application, the first loss and the second loss are processed to obtain the total loss, and the total loss is used to iteratively adjust parameters of the object detection model, so that when adjusting parameters of any network in the position regression network and the classification regression network, the respective parameters can be adjusted based on the total loss that fuses the position regression features and the classification features, thereby ensuring the training effect and improving the prediction accuracy of the target object detection model.
[0114] In an embodiment, the steps (c31-c32) are to iteratively adjust parameters of the object detection model according to the total loss determined by the first loss and the second loss to obtain a target object detection model. In addition, the present application can also iteratively adjust parameters of the position regression network in the object detection model by using the first loss, iteratively adjust parameters of the classification regression network in the object detection model by using the second loss, and then construct the target object detection model based on the adjusted position regression network and classification regression network. The specific training method of the model can be flexibly selected according to actual business conditions, and the present application does not limit it.
[0115] Please refer to Figure 4D , which is an effect diagram of target object detection on a to-be-detected image by using the target detection model. The diagram includes a plurality of pigs, and each pig is marked by a rectangular frame (marker frame) with a rotation angle.
[0116] Based on the above embodiments, the present application has the beneficial effects that by introducing the rotation angle feature in the training process, the model can learn the rotation angle feature of the first target object in the sample image. Compared with the traditional horizontal rectangular frame detection method, the target object detection model trained by the model training method proposed in the present application is beneficial to outputting the position and rotation angle of the target object in the image by the target object detection model, reducing the overlap rate between the detection frames corresponding to a plurality of target objects, and thus improving the accuracy of target object detection.
[0117] The application also proposes generating a plurality of candidate box information in the sample image through a candidate box output network, and performing position regression processing on the candidate box information through a regression output network to obtain a plurality of prediction box information corresponding to the plurality of candidate box information. The difference between the prediction box and the candidate box is that the candidate box information of the candidate box is pre-set, so that the candidate box can only present a plurality of pre-set forms; and the prediction box is obtained based on the position regression processing of the candidate box information. The position regression processing can be regarded as a fine-tuning of the candidate box, so that the prediction box can more accurately and more finely present various forms of the target object (for example, each prediction box can be any combination of any size, any aspect ratio and any rotation angle). Through the above method, the candidate box is position-regressed to obtain the prediction box, so that the prediction box can be more matched with the real position and rotation angle of the target object in the image. Subsequent model training based on the prediction box obtained by position regression of the candidate box can improve the prediction accuracy of the model.
[0118] The application also proposes determining the first prediction box information in the plurality of prediction box information based on the coincidence degree of the candidate box corresponding to the plurality of prediction box information and the labeled box. The number of the first prediction box corresponding to the first prediction box information can be one or more, and the one or more first prediction boxes can indicate the position and rotation angle of the target object in the sample image to a certain extent. Training the object detection model based on the first prediction box improves the model training effect and improves the prediction accuracy of the model. The application also proposes iteratively adjusting the parameters of the object detection model based on the total loss determined by the first loss and the second loss, so that when adjusting the parameters of any network in the position regression network and the classification regression network, the parameters corresponding to each network can be adjusted based on the total loss fused with the position regression feature and the classification feature, thereby ensuring the training effect and improving the prediction accuracy of the target object detection model.
[0119] Please refer to Figure 5 , Figure 5 is a flowchart of an image processing method provided by an exemplary embodiment of the application. The method is applied to a server for detecting a target object (for example, the server 202 in Figure 2 ), and can include the following steps:
[0120] S501, acquiring an image to be detected.
[0121] In an embodiment, the image to be detected is an image in which the second target object needs to be detected by using the target object detection model. The image to be detected can contain at least one second target object, and the second target object contained in the image to be detected can be detected by the method provided in steps S501-S503. The second target object can be a person, an animal (for example, a pig), a plant, or the like in the image. The second target object can be the same as the first target object (for example, a pig), or can be different. The image to be detected can be a top view image of the second target object, or an image in which the second target object is at an arbitrary angle with respect to a reference plane.
[0122] S502, performing object detection processing on the image to be detected by using the target object detection model to determine second bounding box information of the second target object in the image to be detected, the second bounding box information being used to indicate a position of the second target object in the image to be detected and a rotation angle of the second target object with respect to the reference axis.
[0123] In an embodiment of the present application, the second bounding box information of the second target object is used to indicate a position of the second target object in the image to be detected and a rotation angle of the second target object with respect to the reference axis. That is, the second bounding box information can represent a rotation angle of a posture of the second target object in the sample image with respect to the reference axis (for example, a horizontal reference axis or a vertical reference axis based on the sample image frame).
[0124] The target object detection model can be obtained by training the object detection model based on the model training method provided in the foregoing embodiments. For details of the training process, please refer to the related description of steps S301-S304, which will not be repeated here.
[0125] It should be noted that, in the training process of the target detection model, the first prediction box information, the first candidate box information, and the first bounding box information are used for training. The first candidate box information is a candidate box in all candidate boxes generated in the sample image, and the coincidence degree of the candidate box corresponding to the first bounding box information is greater than a preset threshold. That is, in the training process, the first candidate box information and the first bounding box information are positive sample data. In the application process of the target detection model, the second target object is processed based on the positive sample data and the negative sample data to obtain the second bounding box information of the second target object. The process of performing object detection processing on the image to be detected by using the target object detection model to determine the second bounding box information of the second target object in the image to be detected will be introduced as follows:
[0126] In the first step, the target object detection model first extracts features of the to-be-detected image by using a feature extraction network to obtain at least one feature map corresponding to the to-be-detected image. In the second step, for any feature map, the target object detection model finds a plurality of pixel points in the to-be-detected image to which all feature points included in the feature map are mapped. In the third step, for any pixel point, the target object detection model generates a plurality of candidate box information corresponding to the pixel point by using a candidate box output network, and any two pieces of candidate box information in the plurality of candidate box information are different in one or more of a size parameter, an aspect ratio parameter, and a rotation angle parameter. In the fourth step, the target object detection model performs position regression processing on all candidate box information corresponding to the plurality of pixel points in the to-be-detected image by using a position regression network to obtain a plurality of prediction box information. In the fifth step, the target object detection model performs classification regression processing on all candidate box information corresponding to the plurality of pixel points in the to-be-detected image by using a classification regression network to obtain a confidence degree corresponding to each piece of candidate box information. The confidence degree of each piece of candidate box information is the same as the confidence degree of the prediction box information obtained by the position regression processing of the candidate box information. In the sixth step, the target object detection model selects, from the prediction box information corresponding to all candidate box information in the plurality of pixel points in the to-be-detected image, prediction box information whose confidence degree satisfies a condition as to-be-screened prediction box information (the type of the prediction box information is prediction box information of a second target object), and performs non-maximum suppression processing on the to-be-screened prediction box information to finally obtain at least one second marking box information. Each second marking box information is used to generate a marking box indicating a second target object in the to-be-detected image.
[0127] In S503, a marking box is added to the second target object in the to-be-detected image according to the second marking box information.
[0128] The second marking box information includes a center point coordinate (x, y) of a second marking box, a width (w) of the second marking box, a height (h) of the second marking box, and a rotation angle of the second marking box relative to a reference axis. According to the second marking box information, a marking box can be added to the second target object in the to-be-detected image.
[0129] Based on the above embodiments, the application has the beneficial effects that: the application frames the detected second target object by the rotatable marking box, is conducive to the target object detection model to output the position and rotation angle of the second target object in the image, so that the detection result can be more matched with the real position and rotation angle of the second target object in the image, thereby improving the accuracy of target object detection. Compared with the horizontal rectangular box obtained by the traditional target detection algorithm, the rotation angle of the marking box generated by the application can be arbitrary with respect to the reference axis, which reduces the overlap rate between the marking boxes corresponding to multiple second target objects, and also facilitates the operator to more intuitively view the detection result through the marking box, thereby improving the experience.
[0130] Please refer to Figure 6 , Figure 6 is a schematic block diagram of a processing device provided by an embodiment of the application. In an embodiment, the processing device can specifically include:
[0131] The data acquisition module 601 is configured to acquire a sample image and acquire first marking box information corresponding to the sample image, the first marking box information being used to indicate the position of a first target object in the sample image and the rotation angle of the first target object with respect to a reference axis.
[0132] The processing module 602 is configured to input the sample image into the object detection model to obtain a plurality of candidate box information corresponding to the sample image output by a candidate box output network and a plurality of predicted box information corresponding to the plurality of candidate box information output by a regression output network, the candidate box information being used to indicate the candidate position and candidate rotation angle of the target object output by the object detection model, and the predicted box information being obtained by performing position regression on the candidate box information.
[0133] The processing module 602 is further configured to determine first predicted box information from the plurality of predicted box information according to the plurality of candidate box information and the first marking box information, the coincidence degree between the candidate box corresponding to the first candidate box information corresponding to the first predicted box information and the marking box corresponding to the first marking box information being greater than a preset threshold.
[0134] The training module 603 is configured to train the object detection model according to a first relationship parameter between the first predicted box information and the first candidate box information and a second relationship parameter between the first marking box information and the first candidate box information, to obtain a target object detection model.
[0135] Optionally, when acquiring the first marking box information corresponding to the sample image, the data acquisition module 601 is specifically configured to:
[0136] The target object in the sample image is marked with boundary points, and the minimum circumscribed rectangle of the boundary points of the target object is calculated.
[0137] The position parameters of the minimum circumscribed rectangle are determined as the first label box information corresponding to the sample image, and the position parameters of the minimum circumscribed rectangle include the position of the minimum circumscribed rectangle in the sample image and the rotation angle of the minimum circumscribed rectangle relative to the reference axis.
[0138] Optionally, the regression output network includes a position regression network and a classification regression network, the position regression network is configured to output the plurality of prediction box information, and the classification regression network is configured to output the confidence of each of the plurality of prediction box information.
[0139] The training module 603 is configured to train the object detection model according to the first relationship parameter between the first prediction box information and the first candidate box information, and the second relationship parameter between the first label box information and the first candidate box information, to obtain the target object detection model.
[0140] According to the first relationship parameter between the first prediction box information and the first candidate box information, and the second relationship parameter between the first label box information and the first candidate box information, the first loss of the object detection model is calculated, and the first loss is used to indicate the accuracy of the object detection model in positioning the position of the first target object.
[0141] According to the confidence of each of the plurality of prediction box information, the second loss of the object detection model is calculated, and the second loss is used to indicate the accuracy of the object detection model in classification.
[0142] According to the first loss and the second loss, the object detection model is iteratively parameterized to obtain the target object detection model.
[0143] Optionally, when the training module 603 is configured to iteratively parameterize the object detection model according to the first loss and the second loss to obtain the target object detection model, the training module 603 is specifically configured to:
[0144] According to the first loss, the second loss and the number of first prediction box information, the total loss of the object detection model is calculated.
[0145] According to the total loss, the object detection model is iteratively parameterized to obtain the target object detection model.
[0146] Optionally, the processing module 602 is specifically configured for:
[0147] determining, according to the plurality of candidate box information and the first marking box information, an overlap degree between a candidate box corresponding to each of the plurality of candidate box information and a marking box corresponding to the first marking box information;
[0148] determining the candidate box with the overlap degree greater than the preset threshold as a first candidate box;
[0149] determining, as the first prediction box information, prediction box information corresponding to a prediction box corresponding to the first candidate box.
[0150] Optionally, the object detection model further comprises a feature extraction network, the feature extraction network is configured to perform feature extraction on the sample image to obtain a feature map corresponding to the sample image; the plurality of prediction box information comprises a plurality of groups of prediction box information, one group of prediction box information comprises a plurality of prediction box information corresponding to one feature point in one feature map, and any two prediction box information in the plurality of prediction box information corresponding to one feature point are unmatched, the any two prediction box information being unmatched means that one or more of size parameters, aspect ratio parameters and rotation angle parameters of two prediction boxes corresponding to the any two prediction box information are different.
[0151] In an embodiment, the processing apparatus can specifically comprise:
[0152] a data acquisition module 601 configured to acquire a to-be-detected image;
[0153] a processing module 602 configured to perform object detection processing on the to-be-detected image by using a target object detection model to determine second marking box information of a second target object in the to-be-detected image, the second marking box information being used to indicate a position of the second target object in the to-be-detected image and a rotation angle of the second target object relative to a reference axis, the target object detection model being obtained by using the model training method in any one of claims 1-6;
[0154] a marking box output module 604 configured to add a marking box for the second target object in the to-be-detected image according to the second marking box information.
[0155] It should be noted that the functions of the functional modules of the processing apparatus of the embodiments of the present application can be specifically implemented according to the methods in the method embodiments, and the specific implementation process can be referred to the related description of the method embodiments, which will not be described here.
[0156] Please refer to Figure 7 , Figure 7is a schematic block diagram of a computer device provided by an embodiment of the present application. The intelligent terminal in the embodiment shown in the figure can include a processor 701, a storage device 702, and a communication interface 703. The processor 701, the storage device 702, and the communication interface 703 can interact with each other.
[0157] The storage device 702 can include a volatile memory such as a random-access memory (RAM), and can also include a non-volatile memory such as a flash memory, a solid-state drive (SSD), etc. The storage device 702 can also include a combination of the above types of memories.
[0158] The processor 701 can be a central processing unit (CPU). In an embodiment, the processor 701 can also be a graphics processing unit (GPU). The processor 701 can also be a combination of a CPU and a GPU. In an embodiment, the storage device 702 is configured to store program instructions, and the processor 701 can invoke the program instructions to perform the following operations:
[0159] Obtain a sample image, and obtain first bounding box information corresponding to the sample image, the first bounding box information being used to indicate a position of a first target object in the sample image and a rotation angle of the first target object relative to a reference axis;
[0160] Input the sample image into the object detection model to obtain a plurality of candidate box information corresponding to the sample image output by the candidate box output network and a plurality of predicted box information corresponding to the plurality of candidate box information output by the regression output network, the candidate box information being used to indicate a candidate position and a candidate rotation angle of the target object output by the object detection model, and the predicted box information being obtained by performing position regression on the candidate box information;
[0161] According to the plurality of candidate box information and the first bounding box information, determine first predicted box information from the plurality of predicted box information, a candidate box corresponding to first candidate box information corresponding to the first predicted box information having a coincidence degree greater than a preset threshold with a bounding box corresponding to the first bounding box information;
[0162] The object detection model is trained according to the first relationship parameter between the first prediction box information and the first candidate box information and the second relationship parameter between the first marked box information and the first candidate box information, to obtain a target object detection model.
[0163] Optionally, the processor 701 is configured to, when acquiring the first marked box information corresponding to the sample image, specifically configured to:
[0164] boundary points of the target object in the sample image are marked, and a minimum circumscribed rectangle of the boundary points of the target object is calculated;
[0165] a position parameter of the minimum circumscribed rectangle is determined as the first marked box information corresponding to the sample image, the position parameter of the minimum circumscribed rectangle includes a position of the minimum circumscribed rectangle in the sample image and a rotation angle of the minimum circumscribed rectangle relative to a reference axis.
[0166] Optionally, the regression output network includes a position regression network and a classification regression network, the position regression network is configured to output the plurality of prediction box information, and the classification regression network is configured to output a confidence degree corresponding to each of the plurality of prediction box information.
[0167] The processor 701 is configured to, when training the object detection model according to the first relationship parameter between the first prediction box information and the first candidate box information and the second relationship parameter between the first marked box information and the first candidate box information, to obtain a target object detection model, specifically configured to:
[0168] According to the first relationship parameter between the first prediction box information and the first candidate box information and the second relationship parameter between the first marked box information and the first candidate box information, a first loss of the object detection model is calculated, and the first loss is used to indicate the accuracy of the object detection model in positioning the position of the first target object.
[0169] According to the confidence degree corresponding to each of the plurality of prediction box information, a second loss of the object detection model is calculated, and the second loss is used to indicate the accuracy of the object detection model in classification.
[0170] According to the first loss and the second loss, the object detection model is iteratively parameterized to obtain the target object detection model.
[0171] Optionally, the processor 701 is configured to, when iteratively parameterizing the object detection model according to the first loss and the second loss to obtain the target object detection model, specifically configured to:
[0172] calculate a total loss of the object detection model according to the first loss, the second loss, and the number of the first prediction box information;
[0173] perform iterative parameter tuning on the object detection model according to the total loss to obtain the target object detection model.
[0174] Optionally, the processor 701 is specifically configured to:
[0175] determine, according to the plurality of candidate box information and the first marked box information, an overlap degree between a candidate box corresponding to each of the plurality of candidate box information and a marked box corresponding to the first marked box information;
[0176] determine the candidate box with the overlap degree greater than the preset threshold as a first candidate box;
[0177] determine the prediction box information corresponding to the prediction box corresponding to the first candidate box as the first prediction box information.
[0178] Optionally, the object detection model further comprises a feature extraction network, the feature extraction network is configured to perform feature extraction on the sample image to obtain a feature map corresponding to the sample image; the plurality of prediction box information comprises a plurality of groups of prediction box information, one group of prediction box information comprises a plurality of prediction box information corresponding to one feature point in one feature map, and any two prediction box information in the plurality of prediction box information corresponding to one feature point are unmatched, the any two prediction box information being unmatched means that one or more of size parameters, aspect ratio parameters, and rotation angle parameters of two prediction boxes corresponding to the any two prediction box information are different.
[0179] In another embodiment, the storage device 702 is configured to store program instructions, and the processor 701 can invoke the program instructions to perform the following operations:
[0180] obtain a to-be-detected image;
[0181] perform object detection processing on the to-be-detected image by the target object detection model to determine second marked box information of a second target object in the to-be-detected image, the second marked box information being used to indicate a position of the second target object in the to-be-detected image and a rotation angle relative to a reference axis, the target object detection model being obtained by the model training method of any one of claims 1-6;
[0182] add a marked box for the second target object in the to-be-detected image according to the second marked box information.
[0183] In specific implementation, the processor 701, storage device 702, and communication interface 703 described in the embodiments of this application can execute the embodiments of this application. Figure 3 or Figure 5 The implementation methods described in the relevant embodiments of the provided model training methods or image processing methods can also be used to execute the embodiments of this application. Figure 6 The implementation methods described in the relevant embodiments of the provided processing device will not be repeated here.
[0184] In the several embodiments provided in this application, it should be understood that the disclosed methods, apparatuses, and systems can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for example, the division of units is merely a logical functional division, and other division methods may exist in actual implementation; for example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0185] Furthermore, it should be noted that this application embodiment also provides a computer-readable storage medium storing a computer program executed by the aforementioned processing device, and the computer program includes program instructions. When the processor executes the program instructions, it can execute the aforementioned... Figure 3 , Figure 5 The methods described in the corresponding embodiments will not be repeated here. Furthermore, the beneficial effects of using the same methods will also not be repeated. For technical details not disclosed in the computer-readable storage medium embodiments related to this application, please refer to the description of the method embodiments of this application. As an example, program instructions can be deployed on a computer device, or executed on multiple computer devices located in one location, or executed on multiple computer devices distributed across multiple locations and interconnected via a communication network. These multiple computer devices distributed across multiple locations and interconnected via a communication network can constitute a blockchain system.
[0186] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the aforementioned... Figure 3 , Figure 5 The methods described in the corresponding embodiments are therefore not repeated here.
[0187] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The above-mentioned program can be stored in a computer readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiment methods. The storage medium can be a magnetic disc, an optical disc, a read-only memory (ROM), a random access memory (RAM), or the like.
[0188] The above only discloses some embodiments of the present application, and of course cannot limit the scope of the present application. Those skilled in the art can understand that all or part of the processes of the above-mentioned embodiments can be implemented, and equivalent changes made according to the claims of the present application still fall within the scope of the present application.
Claims
1. A model training method, characterized in that, The method is applied to an object detection model, and the object detection model comprises a candidate box output network and a regression output network; the method comprises the following steps: obtaining a sample image and first marking box information corresponding to the sample image, the first marking box information being used to indicate the position of a first target object in the sample image and the rotation angle of the first target object relative to a reference axis; inputting the sample image into the object detection model to obtain a plurality of candidate box information corresponding to the sample image output by the candidate box output network and a plurality of prediction box information corresponding to the plurality of candidate box information output by the regression output network, the candidate box information being used to indicate the candidate position and the candidate rotation angle of the first target object output by the object detection model, and the prediction box information being obtained by performing position regression on the candidate box information; determining first prediction box information in the plurality of prediction box information according to the plurality of candidate box information and the first marking box information, the coincidence degree of the candidate box corresponding to the first candidate box information corresponding to the first prediction box information and the marking box corresponding to the first marking box information being greater than a preset threshold; training the object detection model according to a first relationship parameter between the first prediction box information and the first candidate box information and a second relationship parameter between the first marking box information and the first candidate box information to obtain a target object detection model.
2. The method of claim 1, wherein, The first marking box information corresponding to the sample image is obtained by the following steps: performing boundary point marking on the first target object in the sample image and calculating the minimum circumscribed rectangle of the boundary points of the first target object; determining the position parameter of the minimum circumscribed rectangle as the first marking box information corresponding to the sample image, the position parameter of the minimum circumscribed rectangle comprising the position of the minimum circumscribed rectangle in the sample image and the rotation angle of the minimum circumscribed rectangle relative to the reference axis.
3. The method of claim 1, wherein, The regression output network comprises a position regression network and a classification regression network, the position regression network being used to output the plurality of prediction box information, and the classification regression network being used to output the confidence degree corresponding to each of the plurality of prediction box information; The object detection model is trained according to the first relationship parameter between the first prediction box information and the first candidate box information and the second relationship parameter between the first marking box information and the first candidate box information to obtain a target object detection model, and the method comprises the following steps: calculating a first loss of the object detection model according to the first relationship parameter between the first prediction box information and the first candidate box information and the second relationship parameter between the first marking box information and the first candidate box information, the first loss being used to indicate the accuracy of the object detection model in positioning the position of the first target object; calculating a second loss of the object detection model according to the confidence degree corresponding to each of the plurality of prediction box information, the second loss being used to indicate the classification accuracy of the object detection model; iteratively adjusting the parameters of the object detection model according to the first loss and the second loss to obtain the target object detection model.
4. The method of claim 3, wherein, The iterative parameter adjustment is performed on the object detection model according to the first loss and the second loss to obtain the target object detection model, and the method comprises the steps of: calculating a total loss of the object detection model according to the first loss, the second loss and the number of the first prediction box information; performing iterative parameter adjustment on the object detection model according to the total loss to obtain the target object detection model.
5. The method according to any one of claims 1-4, characterized in that, The first prediction box information is determined from the plurality of prediction box information according to the plurality of candidate box information and the first marking box information, and the method comprises the steps of: determining the coincidence degree between the candidate box corresponding to each of the plurality of candidate box information and the marking box corresponding to the first marking box information according to the plurality of candidate box information and the first marking box information; determining the candidate box with a coincidence degree greater than the preset threshold as the first candidate box; determining the prediction box information corresponding to the prediction box corresponding to the first candidate box as the first prediction box information.
6. The method according to any one of claims 1-4, characterized in that, The object detection model further comprises a feature extraction network, the feature extraction network is used for feature extraction on the sample image to obtain a feature map corresponding to the sample image; the plurality of prediction box information comprises a plurality of groups of prediction box information, one group of prediction box information comprises a plurality of prediction box information corresponding to one feature point in one feature map, and any two prediction box information corresponding to one feature point are not matched, which means that the size parameter, the aspect ratio parameter and the rotation angle parameter of the two prediction boxes corresponding to the any two prediction box information are different.
7. An image processing method characterized by, The method comprises: obtaining a to-be-detected image; performing object detection processing on the to-be-detected image by a target object detection model to determine second marking box information of a second target object in the to-be-detected image, the second marking box information being used to indicate the position of the second target object in the to-be-detected image and the rotation angle relative to a reference axis, the target object detection model being obtained by the model training method in any one of claims 1-6; adding a marking box for the second target object in the to-be-detected image according to the second marking box information.
8. A processing device, characterized by The device comprises a module for implementing the model training method in any one of claims 1-6, or a module for implementing the image processing method in claim 7.
9. A computer device, comprising: comprises: a processor, a storage device and a communication interface, which are connected to each other, wherein the storage device stores executable program codes, and the processor is used to invoke the executable program codes to implement the model training method in any one of claims 1-6, or a module for implementing the image processing method in claim 7.
10. A computer readable storage medium characterized by, The computer readable storage medium stores a computer program, the computer program comprises program instructions, and the program instructions are executed by a processor to implement the model training method in any one of claims 1-6, or a module for implementing the image processing method in claim 7.
Citation Information
Patent Citations
Target detection model training method and device and electronic equipment
CN111738072A
Target detection and model training method and device, computer equipment and storage medium
CN114694218A