Detection model training method, object detection method and electronic equipment

By setting different preset thresholds for detection targets and obstacles and adjusting detection model parameters, the accuracy problem of electronic devices when identifying detection targets is solved, achieving higher recognition accuracy and user experience.

CN120388249APending Publication Date: 2025-07-29UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510388146.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

In the prior art, electronic devices have poor accuracy when identifying and detecting targets, which affects normal operation.

Method used

By setting different preset thresholds for detection targets and obstacles, and adjusting the model parameters of the detection model, the detection model can accurately distinguish detection targets and obstacles, and improving identification accuracy.

Benefits of technology

It improves the accuracy of the detection model to identify detection targets, and improves the object detection capabilities and user experience of electronic devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388249A_ABST
    Figure CN120388249A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of image processing, and particularly relates to a training method of a detection model, an object detection method and electronic equipment. In the method, when a detection model is trained, a common obstacle can be trained through second labeling information and a second preset threshold value, a detection target can be trained through first labeling information and a first preset threshold value, and the first preset threshold value is different from the second preset threshold value. The detection target and the common obstacle can be trained based on the first preset threshold value and the second preset threshold value which are different, so that the detection model obtained through training can accurately distinguish the detection target and the obstacle, the accuracy of the detection model for identifying the detection target can be improved, and the detection efficiency is improved. The object detection accuracy of the electronic equipment is improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of image processing, and particularly relates to a training method for a detection model, an object detection method, and an electronic device. Background Art

[0002] With the development of artificial intelligence (AI) technology, the functions of electronic devices have become more and more powerful. For example, an electronic device can identify a detection target (such as furniture) in the environment based on AI technology (such as visual recognition technology), so as to perform subsequent tasks based on the identified detection target, such as map construction (such as indoor partition mapping), autonomous navigation, cleaning planning, or path optimization. When the accuracy of the electronic device in identifying the detection target is poor, it will affect the normal operation of the electronic device. Therefore, how to improve the accuracy of the electronic device in identifying the detection target has become an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0003] Embodiments of this application provide a training method for a detection model, an object detection method, and an electronic device, which can improve the accuracy of the electronic device in identifying a detection target and enhance the user experience.

[0004] In a first aspect, embodiments of this application provide a training method for a detection model, and the method includes:

[0005] Obtain a training image set; the training image set includes multiple training images, and each training image includes first annotation information and / or second annotation information, the first annotation information includes a first detection frame corresponding to a detection target in the training image, and the second annotation information includes a second detection frame corresponding to an obstacle in the training image;

[0006] Input the training images in the training image set into a detection model for processing, to obtain a third detection frame and a first confidence level corresponding to the detection target in the training image output by the detection model, and / or a fourth detection frame and a second confidence level corresponding to the obstacle in the training image;

[0007] When it is determined that the first confidence level corresponding to the detection target is greater than the first preset threshold, and / or it is determined that the second confidence level corresponding to the obstacle is greater than the second preset threshold, based on the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle, the model parameters of the detection model are adjusted, and the step of inputting the training image in the training image set into the detection model for processing is returned to obtain the third detection box and the first confidence level corresponding to the detection target in the training image output by the detection model, and / or the fourth detection box and the second confidence level corresponding to the obstacle in the training image, and subsequent steps are performed until the detection model meets the preset conditions, and a trained detection model is obtained. The first preset threshold is different from the second preset threshold.

[0008] In the training method of the detection model provided above, when training the detection model, for the detection target and ordinary obstacles, different first preset thresholds and second preset thresholds can be used for training processing, that is, the model parameters of the detection model can be adjusted based on different preset thresholds, so that the trained detection model can accurately distinguish the detection target and the obstacle, thereby improving the accuracy of the detection model in identifying the detection target (such as furniture), improving the accuracy of the electronic device in object detection, and enhancing the user experience.

[0009] In some embodiments, the adjusting the model parameters of the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle includes:

[0010] Determine the detection accuracy corresponding to the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle;

[0011] When it is determined that the detection accuracy is less than or equal to the third preset threshold, adjust the model parameters of the detection model.

[0012] In some embodiments, the first detection box is the detection box corresponding to the detection target that is completely displayed in the training image; the second detection box is the detection box corresponding to the obstacle that is completely displayed or partially displayed in the training image.

[0013] In some embodiments, the first preset threshold is greater than the second preset threshold.

[0014] In a second aspect, an embodiment of the present application provides an object detection method, and the method includes:

[0015] Obtain a first image corresponding to the target area;

[0016] Input the first image into the trained detection model for processing to obtain the detection result output by the detection model;

[0017] Wherein, the detection model is trained based on the training method of the detection model provided in the above first aspect.

[0018] In some embodiments, the detection result includes the categories corresponding to one or more detected objects, the detection frames corresponding to each detected object, and the confidence levels corresponding to each detected object.

[0019] In one embodiment, after inputting the first image into the trained detection model for processing to obtain the detection result output by the detection model, the method further includes:

[0020] When it is determined that the detected object includes a detection target according to the category corresponding to the detected object, obtain the first confidence level corresponding to the detection target;

[0021] When it is determined that the first confidence level is greater than the first preset threshold, retain the detection frame corresponding to the detection target.

[0022] In another embodiment, after inputting the first image into the trained detection model for processing to obtain the detection result output by the detection model, the method further includes:

[0023] When it is determined that the detected object includes an obstacle according to the category corresponding to the detected object, obtain the second confidence level corresponding to the obstacle;

[0024] When it is determined that the second confidence level is greater than the second preset threshold, retain the detection frame corresponding to the obstacle.

[0025] Exemplarily, the first preset threshold is greater than the second preset threshold.

[0026] In a third aspect, an embodiment of the present application provides a training device for a detection model, including:

[0027] An image acquisition module, configured to acquire a training image set; the training image set includes multiple training images, and each training image includes first annotation information and / or second annotation information, the first annotation information includes a first detection frame corresponding to a detection target in the training image, and the second annotation information includes a second detection frame corresponding to an obstacle in the training image;

[0028] A processing module, configured to input the training images in the training image set into a detection model for processing, to obtain a third detection box and a first confidence level corresponding to the detection target in the training images output by the detection model, and / or a fourth detection box and a second confidence level corresponding to the obstacle in the training images;

[0029] An adjustment module, configured to, when it is determined that the first confidence level corresponding to the detection target is greater than a first preset threshold, and / or it is determined that the second confidence level corresponding to the obstacle is greater than a second preset threshold, adjust the model parameters of the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle, and return to the execution of the processing module until the detection model meets a preset condition, to obtain a trained detection model, where the first preset threshold is different from the second preset threshold.

[0030] In a fourth aspect, an embodiment of the present application provides an object detection device, including:

[0031] An image acquisition module, configured to acquire a first image corresponding to a target area;

[0032] An object detection module, configured to input the first image into a trained detection model for processing, to obtain a detection result output by the detection model;

[0033] Wherein, the detection model is trained based on the detection model training method in the first aspect above.

[0034] In a fifth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the detection model training method in any one of the first aspects above, or the electronic device implements the object detection method in any one of the second aspects above.

[0035] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program. When the computer program is executed by an electronic device, the electronic device implements the detection model training method in any one of the first aspects above, or the electronic device implements the object detection method in any one of the second aspects above.

[0036] In a seventh aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by an electronic device, it enables the electronic device to implement the training method of the detection model described in any one of the above first aspects, or enables the electronic device to implement the object detection method described in any one of the above second aspects.

[0037] It can be understood that for the beneficial effects of the above second aspect to seventh aspect, reference can be made to the relevant descriptions in the above first aspect, and details are not repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0039] Figure 1 It is a schematic flowchart of a method for training a detection model provided by an embodiment of the present application;

[0040] Figure 2 It is a schematic flowchart of an object detection method provided by an embodiment of the present application;

[0041] Figure 3 It is a schematic structural diagram of a device for training a detection model provided by an embodiment of the present application;

[0042] Figure 4 It is a schematic structural diagram of an object detection device provided by an embodiment of the present application;

[0043] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system structures and technologies are presented in order to thoroughly understand the embodiments of the present application. However, those skilled in the art should understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0045] It should be understood that when used in the specification of the present application and the appended claims, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0046] It should also be understood that the term "and / or" as used in the specification of this application and the appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0047] As used in the specification of this application and the appended claims, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" depending on the context. Similarly, the phrases "if determined" or "if [the described condition or event] is detected" can be interpreted as meaning "once determined", "in response to determining", "once [the described condition or event] is detected", or "in response to detecting [the described condition or event]" depending on the context.

[0048] In addition, in the description of the specification of this application and the appended claims, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and should not be construed as indicating or implying relative importance.

[0049] Reference to "one embodiment" or "some embodiments" etc. described in the specification of this application means that a particular feature, structure, or characteristic described in connection with the embodiment is included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having", and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0050] In the recognition of indoor scenes, an electronic device generally needs to recognize various detection targets (such as furniture like beds, sofas, dining tables, or TVs, etc.) to construct an indoor partition map based on the distribution of the detection targets. For example, the partition where the bed is located can generally be marked as the bedroom, and the partitions where the sofa, dining table, TV, etc. are located can generally be marked as the living room, and so on. Among them, the electronic device generally uses a trained detection model to recognize objects, and the detection model is generally trained based on training images with prior object annotation. For example, the detection model can be trained with training images where object annotation is performed based on a general annotation strategy (that is, as long as an object appears in the image, regardless of whether the object is complete, a detection box will be annotated). This training method is likely to mislead the detection model to learn local information of the detection target (such as furniture) and ignore the overall features of the detection target (such as furniture) as a large structural object, resulting in misrecognition of the detection target by the electronic device and poor accuracy of the electronic device in recognizing the detection target. When the accuracy of the electronic device in recognizing the detection target is poor, it will affect the normal operation of the electronic device. Therefore, how to improve the accuracy of the electronic device in detecting target recognition has become an urgent problem for those skilled in the art to solve.

[0051] To solve the above problems, embodiments of the present application provide a method for training a detection model, an object detection method, and an electronic device. In the method for training a detection model, an electronic device can obtain a training image set. Each training image in the training image set can include first annotation information and / or second annotation information. The first annotation information can include a first detection box corresponding to a detection target in the training image, and the second annotation information can include a second detection box corresponding to an obstacle in the training image. Subsequently, the electronic device can input the training images in the training image set into the detection model for processing, and obtain a third detection box and a first confidence level corresponding to the detection target in the training image output by the detection model, and / or a fourth detection box and a second confidence level corresponding to the obstacle in the training image. When it is determined that the first confidence level corresponding to the detection target is greater than a first preset threshold, and / or it is determined that the second confidence level corresponding to the obstacle is greater than a second preset threshold, the electronic device can adjust the model parameters of the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle, and continue training based on the detection model after the model parameters are adjusted until the detection model meets the preset conditions. Among them, the first preset threshold is different from the second preset threshold. That is, in the embodiments of the present application, when training the detection model, for the detection target and ordinary obstacles, different first preset thresholds and second preset thresholds can be used for training processing, that is, the model parameters of the detection model can be adjusted based on different preset thresholds, so that the trained detection model can accurately distinguish the detection target and the obstacle, thereby improving the accuracy of the detection model in recognizing the detection target, improving the accuracy of the electronic device in object detection, enhancing the user experience, and having strong usability and practicality.

[0052] The method for training a detection model provided by the embodiments of the present application can be applied to electronic devices such as robots (such as sweeping robots, handling robots, or guiding robots), mobile phones, tablet computers, wearable devices, vehicle-mounted devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), or cloud servers. The object detection method provided by the embodiments of the present application can also be applied to electronic devices such as robots (such as sweeping robots, handling robots, or guiding robots), mobile phones, tablet computers, wearable devices, vehicle-mounted devices, AR / VR devices, laptop computers, UMPCs, netbooks, or PDAs. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices to which the method for training a detection model or the object detection method is applied.

[0053] It should be noted that the electronic device used for the training method of the detection model and the object detection method can be the same electronic device or different electronic devices, and the embodiments of the present application do not limit this. For example, it can be determined according to the actual scenario that the electronic devices used for the training method of the detection model and the object detection method are both floor sweeping robots. For example, it can be determined according to the actual scenario that the electronic device used for the training method of the detection model is a cloud server, and the electronic device used for the object detection method is a floor sweeping robot. That is, after the detection model is trained on the cloud server, the cloud server can send the detection model to the floor sweeping robot. The floor sweeping robot can perform object detection based on the trained detection model.

[0054] Next, the training method of the detection model provided by the embodiments of the present application will be described in detail in conjunction with the accompanying drawings and specific application scenarios.

[0055] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of the training method of the detection model provided by the embodiments of the present application. This method can be applied to an electronic device. As Figure 1 shown, this method can include:

[0056] S101. The electronic device obtains a training image set; the training image set includes multiple training images, and each training image includes first annotation information and / or second annotation information. The first annotation information includes a first detection box corresponding to a detection target in the training image, and the second annotation information includes a second detection box corresponding to an obstacle in the training image.

[0057] It should be noted that the detection target can be an object of a furniture category such as a bed, a sofa, a dining table, or a TV. The obstacle can be an object of an obstacle category such as a shoe, a ball of thread, or a fabric. That is, the detection target can be an object of a furniture category associated with the spatial division of the environment. The obstacle can be an object of an obstacle category other than the object of the furniture category that needs to be bypassed during the movement of the electronic device. It should be understood that the object of the obstacle category in the embodiments of the present application may not include the object of the furniture category.

[0058] Exemplarily, the annotation information (such as the first annotation information and the second annotation information) can include information such as the category of the object and the area corresponding to the object (that is, the area of the object in the image, which can also be called the detection box). Among them, the category of the object corresponding to the object can include the furniture category and the obstacle category. In addition, the annotation information can also include the specific type of the object, for example, it can include a bed, a sofa, a dining table, a TV, a shoe, a ball of thread, or a fabric, etc.

[0059] In some embodiments, different annotation methods can be used to annotate objects of the furniture category and objects of the obstacle category. Exemplarily, for objects of the obstacle category (i.e., obstacles), a general annotation method can be used (that is, as long as the object appears in the image, regardless of whether the object is complete, the category corresponding to the object and the region corresponding to the object (such as the second detection box) can be annotated, etc.) to obtain the second annotation information corresponding to the obstacle. For objects of the furniture category (i.e., detection targets), a weak annotation method can be used (that is, only when the object appears completely in the image, the category corresponding to the object and the region corresponding to the object (such as the first detection box) will be annotated, etc.) to obtain the first annotation information corresponding to the detection target.

[0060] In the embodiments of the present application, by using the weak annotation method to annotate objects of the furniture category (i.e., detection targets), it is only necessary to annotate the complete part of the furniture object, rather than annotating each part of the furniture object. This can not only save the annotation cost, but also reduce the misannotation of local information (for example, the table legs are misannotated as the table body), and can guide the detection model to learn the overall features of the furniture object, so that the detection model can accurately learn the overall features of the furniture object, and thus the detection model can accurately identify the furniture object, improving the accuracy of the detection model in identifying the furniture object.

[0061] S102. The electronic device inputs the training images in the training image set into the detection model for processing, and obtains the third detection box and the first confidence level corresponding to the detection target in the training images output by the detection model, and / or the fourth detection box and the second confidence level corresponding to the obstacle in the training images.

[0062] In the embodiments of the present application, after obtaining the training image set, the electronic device can train the detection model through the training image set, that is, the electronic device can input each training image in the training image set into the detection model for processing, and obtain the third detection box and the first confidence level corresponding to the detection target in each training image output by the detection model, and / or the fourth detection box and the second confidence level corresponding to the obstacle in each training image.

[0063] In some embodiments, the detection model can be a detection model based on the YOLOv6n network.

[0064] It should be noted that the above-mentioned detection model being a detection model based on the YOLOv6n network is only for exemplary explanation and should not be construed as a limitation on the embodiments of the present application. In the embodiments of the present application, the detection model can also be a detection model based on other network structures.

[0065] S103. The electronic device determines that the first confidence level corresponding to the detection target is greater than the first preset threshold, and / or determines that the second confidence level corresponding to the obstacle is greater than the second preset threshold.

[0066] S104. The electronic device determines whether the detection model meets the preset conditions according to the first detection frame and the third detection frame corresponding to the detection target, and / or the second detection frame and the fourth detection frame corresponding to the obstacle.

[0067] In some embodiments, when training the detection model, the electronic device can process the detection task corresponding to the detection target according to the first preset threshold, and can process the detection task corresponding to the obstacle according to the second preset threshold, so as to adjust the model parameters of the detection model according to the processing results.

[0068] That is, after inputting each training image into the detection model for processing to obtain the third detection frame and the first confidence level corresponding to the detection target in each training image output by the detection model, and / or the fourth detection frame and the second confidence level corresponding to the obstacle in each training image, the electronic device can process the detection results according to the first preset threshold and the second preset threshold, so as to determine the final detection result of the detection model based on the processing results, and thus adjust the model parameters of the detection model based on the final detection result of the detection model and the first annotation information and / or the second annotation information corresponding to each training image. For example, the electronic device can determine the detection target actually detected by the detection model according to the first preset threshold and the first confidence level (for example, the detection target with the first confidence level greater than the first preset threshold), and can determine the obstacle actually detected by the detection model according to the second preset threshold and the second confidence level (for example, the obstacle with the second confidence level greater than the second preset threshold), so as to determine whether to adjust the model parameters of the detection model and whether to continue training the detection model according to the detection target and / or the obstacle actually detected by the detection model.

[0069] Exemplarily, for each training image, the detection result output by the detection model may include n detection regions (which may also be referred to as detection frames), and the parameters corresponding to each detection region include: category, coordinates of the detection region, confidence level p, etc. Among them, the category may refer to the category of the object corresponding to the detection region. The confidence level p may refer to the degree of certainty of the detection model for the detection region. p ∈ [0, 1], that is, p may be a value between 0 and 1. The larger p is, the greater the confidence level is, and the greater the confidence level is, the more certain the detection model is about the detection region. It should be understood that n ≥ 0.

[0070] In one embodiment, for the obstacle detection task, that is, for the obstacles included in the detection result, the electronic device can process based on a second preset threshold. Exemplarily, when it is determined that the detection result corresponding to a certain training image includes one or more obstacles, the electronic device can respectively obtain the confidence levels (such as the second confidence level) corresponding to each obstacle, and can process the obstacles according to the confidence levels corresponding to each obstacle and the second preset threshold. When it is determined that the confidence level corresponding to a certain obstacle (such as obstacle A) is greater than the second preset threshold, the electronic device can determine that obstacle A is indeed detected in the training image. When it is determined that the confidence level corresponding to obstacle A is less than or equal to the second preset threshold, the electronic device can determine that obstacle A is not detected in the training image.

[0071] In another embodiment, for the detection task of a detection target (such as furniture), that is, for the detection targets included in the detection result, the electronic device can process based on a first preset threshold. Exemplarily, when it is determined that the detection result corresponding to a certain training image includes one or more detection targets, the electronic device can respectively obtain the confidence levels (such as the first confidence level) corresponding to each detection target, and can process the detection targets according to the confidence levels corresponding to each detection target and the first preset threshold. When it is determined that the confidence level corresponding to a certain detection target (such as detection target B) is greater than the first preset threshold, the electronic device can determine that detection target B is indeed detected in the training image. When it is determined that the confidence level corresponding to detection target B is less than or equal to the first preset threshold, the electronic device can determine that detection target B is not detected in the training image.

[0072] In some embodiments, the preset condition can be that the detection accuracy corresponding to the detection model is greater than or equal to a certain preset threshold (such as a third preset threshold that can be called). Among them, the third preset threshold can be determined according to the actual scenario, and the embodiments of the present application do not limit this.

[0073] In the embodiments of the present application, after determining the detection targets and / or obstacles actually detected by the detection model, the electronic device may determine whether the detection model meets the preset conditions according to the first detection frames (i.e., the detection frames marked in the first annotation information) and the third detection frames (i.e., the detection frames detected by the detection model) corresponding to the detection targets (hereinafter referred to as detection target T for easy understanding) actually detected by the detection model, and / or the second detection frames (i.e., the detection frames marked in the second annotation information) and the fourth detection frames (i.e., the detection frames detected by the detection model) corresponding to the obstacles (hereinafter referred to as obstacle O for easy understanding) actually detected by the detection model. That is, the electronic device may determine whether the detection of each detection target T by the detection model is accurate according to the first detection frame and the third detection frame corresponding to each detection target T, and may determine whether the detection of each obstacle O by the detection model is accurate according to the second detection frame and the fourth detection frame corresponding to each obstacle O, so as to determine the detection accuracy corresponding to the detection model. After determining the detection accuracy corresponding to the detection model, the electronic device may determine whether the detection accuracy corresponding to the detection model is greater than or equal to a third preset threshold to determine whether the detection model meets the preset conditions. Among them, when the detection accuracy is greater than or equal to the third preset threshold, it may be determined that the detection model meets the preset conditions. When the detection accuracy is less than the third preset threshold, it may be determined that the detection model does not meet the preset conditions.

[0074] It should be noted that the detection accuracy corresponding to the detection model may be the ratio between the sum of the number of accurately detected detection targets T (for example, S11) and the number of accurately detected obstacles O (for example, S12), and the sum of the actual number of detection targets (for example, S21) and the actual number of obstacles (for example, S22), for example, the detection accuracy corresponding to the detection model = (S11 + S12) / (S21 + S22). Among them, the actual number of detection targets may be the total number of detection targets included in the training image set. The actual number of obstacles may be the total number of obstacles included in the training image set.

[0075] It should be understood that the detection accuracy corresponding to the detection model is the ratio between the sum of the number of accurately detected detection targets T and the number of accurately detected obstacles O and the sum of the actual number of detection targets and the actual number of obstacles corresponding thereto. The preset condition is that the detection accuracy corresponding to the detection model is greater than or equal to a third preset threshold. This is only for exemplary explanation and should not be construed as a limitation on the embodiments of the present application. In the embodiments of the present application, the detection accuracy corresponding to the detection model may include a first detection accuracy and a second detection accuracy. The first detection accuracy may be the detection accuracy corresponding to the detection target, that is, the ratio between the number of accurately detected detection targets T and the actual number of detection targets corresponding thereto. The second detection accuracy may be the detection accuracy corresponding to the obstacle, that is, the ratio between the number of accurately detected obstacles O and the actual number of obstacles corresponding thereto. The preset condition may include that the first detection accuracy is greater than or equal to a certain preset threshold (for example, a fourth preset threshold), and the second detection accuracy is greater than or equal to a certain preset threshold (for example, a fifth preset threshold). Among them, the fourth preset threshold and the fifth preset threshold may be specifically determined according to the actual scenario.

[0076] That is to say, the electronic device can determine whether the detection model accurately detects each detection target T according to the first detection frame and the third detection frame corresponding to each detection target T actually detected by the detection model, so as to determine the first detection accuracy corresponding to the detection model according to the number of accurately detected detection targets T and the actual number of detection targets corresponding thereto. In addition, the electronic device can also determine whether the detection model accurately detects each obstacle O according to the second detection frame and the fourth detection frame corresponding to each obstacle O actually detected by the detection model, so as to determine the second detection accuracy corresponding to the detection model according to the number of accurately detected obstacles O and the actual number of obstacles corresponding thereto. After determining the first detection accuracy and the second detection accuracy corresponding to the detection model, the electronic device can determine whether the first detection accuracy corresponding to the detection model is greater than or equal to the fourth preset threshold, and determine whether the second detection accuracy corresponding to the detection model is greater than or equal to the fifth preset threshold, so as to determine whether the detection model meets the preset condition. Among them, when the first detection accuracy is greater than or equal to the fourth preset threshold and the second detection accuracy is greater than or equal to the fifth preset threshold, it can be determined that the detection model meets the preset condition. When the first detection accuracy is less than the fourth preset threshold, or the second detection accuracy is less than the fifth preset threshold, it can be determined that the detection model does not meet the preset condition.

[0077] Exemplarily, in order to balance the recall rate and precision of obstacle detection, reduce the possibility of missing obstacles and misdetecting non-obstacles, and improve the efficiency of obstacle detection, the second preset threshold can be determined as a relatively balanced value. For example, the second preset threshold can be determined to be 0.5.

[0078] Exemplarily, since the detection target (such as furniture) does not need to be recalled in real time, for an indoor scene to be recognized, it is only necessary to recognize the corresponding furniture once. For example, during the mapping process, it is only necessary to recognize a bed in a certain partition during a certain recognition, and then this partition can be matched as a bedroom. Therefore, the first preset threshold can be determined as a relatively high value to reduce possible false detections during the furniture recognition process, and the reduction in the recall rate caused by the relatively high first preset threshold will not affect the realization of indoor scene recognition. For example, the first preset threshold can be determined as 0.8.

[0079] It should be noted that the determination of the first preset threshold as 0.8 and the second preset threshold as 0.5 described above are only for exemplary explanation and should not be construed as a limitation on the embodiments of the present application. In the implementation of the present application, the first preset threshold can be determined as 0.9 and the second preset threshold as 0.5 according to the actual scenario. For example, the first preset threshold can be determined as 0.9 and the second preset threshold as 0.6 according to the actual scenario, and so on.

[0080] It should be understood that after the first preset threshold is determined, it can also be adjusted in real time according to the actual scenario. Similarly, after the second preset threshold is determined, it can also be adjusted in real time according to the actual scenario.

[0081] S105. When the electronic device determines that the detection model does not meet the preset conditions, it adjusts the model parameters of the detection model and returns to execute the step of inputting the training images in the training image set into the detection model for processing to obtain the third detection box and the first confidence level corresponding to the detection target in the training image, and / or the fourth detection box and the second confidence level corresponding to the obstacle in the training image, and the subsequent steps.

[0082] In the embodiments of the present application, when it is determined that the detection model does not meet the preset conditions, the electronic device can determine that the detection model has not been trained completely. At this time, the electronic device can adjust the model parameters of the detection model and can continue to train the detection model according to the training image set.

[0083] S106. When the electronic device determines that the detection model meets the preset conditions, it obtains the trained detection model.

[0084] When it is determined that the detection model meets the preset conditions, the electronic device can determine that the detection model has been trained completely, and thus obtains the trained detection model.

[0085] In the embodiments of the present application, when training a detection model, for a detection target and ordinary obstacles, different first preset thresholds and second preset thresholds can be used for training processing, that is, the model parameters of the detection model can be adjusted based on different preset thresholds, so that the trained detection model can accurately distinguish the detection target from the obstacles, thereby improving the accuracy of the detection model in identifying the detection target (such as furniture), improving the accuracy of the electronic device in object detection, and enhancing the user experience.

[0086] Next, the object detection method provided by the embodiments of the present application will be described in detail in conjunction with the accompanying drawings and specific application scenarios.

[0087] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of the object detection method provided by the embodiments of the present application. This method can be applied to an electronic device. Hereinafter, taking the application of this method to a sweeping robot as an example for exemplary illustration. As Figure 2 shown, the method may include:

[0088] S201. The sweeping robot acquires a first image corresponding to a target area.

[0089] It should be noted that the target area can be any area where the sweeping robot needs to work. For example, the target area can be areas such as a house, a ward, or an office. Among them, the target area can include all rooms in the entire house, or can also include some rooms in the entire house. The embodiments of the present application do not make specific limitations on this. Similarly, for areas such as a ward or an office, the target area can also be all wards or all office areas on one floor, or can also be some wards or some office areas on one floor.

[0090] In some embodiments, an image acquisition device such as a camera is provided on the sweeping robot. The sweeping robot can collect an image corresponding to the target area (i.e., the first image) through the image acquisition device such as a camera, so as to detect objects in the target area based on the first image.

[0091] S202. The sweeping robot inputs the first image into the trained detection model for processing, and obtains the detection result output by the detection model.

[0092] In some embodiments, a trained detection model can be deployed on the sweeping robot. After acquiring the first image corresponding to the target area, the sweeping robot can process the first image through the trained detection model to determine the detection result corresponding to the target area. For example, the sweeping robot can input the first image into the trained detection model. After the trained detection model acquires the first image, it can process the first image to obtain the detection result corresponding to the target area.

[0093] Exemplarily, the detection model can be a detection model based on the YOLOv6n network. Herein, the detection model being a detection model based on the YOLOv6n network is only for exemplary explanation and should not be construed as a limitation on the embodiments of the present application. The embodiments of the present application can determine the network structure corresponding to the detection model according to the actual scenario. It should be understood that the training process of the detection model can refer to the relevant content shown above. For the sake of brevity, it will not be elaborated herein. Figure 1 As shown above, for the sake of brevity, it will not be elaborated herein.

[0094] Exemplarily, the detection results output by the detection model can include the categories of one or more detected objects, the detection regions corresponding to each detected object, and the confidence levels corresponding to each detected object. Herein, the detected object can refer to the object recognized by the detection model in the first image.

[0095] In some embodiments, after obtaining the detection results, the sweeping robot can perform post-processing on the detection results to determine the final detection results according to the post-processing results. Exemplarily, the sweeping robot can perform different post-processing on the objects of the obstacle category (i.e., obstacles) and the objects of the furniture category (i.e., detection targets) included in the detection results to determine the final detection results.

[0096] In one embodiment, for the detection task of obstacles, that is, for the obstacles included in the detection results, the sweeping robot can perform post-processing based on a second preset threshold. Exemplarily, when it is determined that the detection results corresponding to the first image include one or more obstacles, that is, when it is determined that the detected objects include one or more obstacles, the sweeping robot can respectively obtain the second confidence levels corresponding to each obstacle and can perform post-processing on the obstacles according to the second confidence levels corresponding to each obstacle and the second preset threshold. When it is determined that the second confidence level corresponding to a certain obstacle (e.g., obstacle C) is greater than the second preset threshold, the sweeping robot can determine that obstacle C is indeed detected in the target area, can retain the detection frame corresponding to obstacle C, or can output the detection frame corresponding to obstacle C. When it is determined that the second confidence level corresponding to obstacle C is less than or equal to the second preset threshold, the sweeping robot can consider that obstacle C is not detected in the target area and can not retain the detection frame corresponding to obstacle C.

[0097] In another embodiment, for the detection task of furniture, that is, for the detection targets included in the detection result, the sweeping robot can perform post-processing based on a first preset threshold. Exemplarily, when it is determined that the detection result corresponding to the first image includes one or more detection targets, that is, when it is determined that the detected object includes one or more detection targets, the sweeping robot can respectively obtain the first confidence corresponding to each detection target, and can perform post-processing of the detection target according to the first confidence corresponding to each detection target and the first preset threshold. When it is determined that the first confidence corresponding to a certain detection target (for example, detection target D) is greater than the first preset threshold, the sweeping robot can determine that detection target D is indeed detected in the target area, can retain the detection frame corresponding to detection target D, or can output the detection frame corresponding to detection target D. When it is determined that the first confidence corresponding to detection target D is less than or equal to the first preset threshold, the sweeping robot can consider that detection target D is not detected in the target area and can not retain the detection frame corresponding to detection target D.

[0098] It should be noted that the first preset threshold can be greater than the second preset threshold. Among them, the specific contents of the first preset threshold and the second preset threshold can refer to the relevant descriptions of the foregoing first preset threshold and the second preset threshold. For the sake of simplicity, they will not be elaborated here.

[0099] Exemplarily, in order to balance the recall rate and precision of obstacle detection, reduce the possibility of missing obstacles and misdetecting non-obstacles, improve the accuracy and efficiency of obstacle detection, and enable the sweeping robot to perform efficient and accurate detection of obstacles, the second preset threshold can be determined as a relatively balanced value. For example, the second preset threshold can be determined as 0.5.

[0100] Exemplarily, since furniture does not need to be recalled in real time, for an indoor scene to be recognized, it is only necessary to recognize the corresponding furniture once. For example, during the mapping process, as long as the bed is recognized in a certain partition during a certain recognition, that partition can be matched as a bedroom. Therefore, the first preset threshold can be determined as a relatively high value to reduce the possible misdetection during the furniture recognition process, and the reduction in the recall rate caused by the relatively high first preset threshold will not affect the realization of indoor scene recognition. For example, the first preset threshold can be determined as 0.8.

[0101] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0102] Corresponding to the training method of the detection model described in the foregoing embodiments, Figure 3The structural block diagram of the training device for the detection model provided by the embodiments of the present application is shown. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.

[0103] Referring to Figure 3 , the device may include:

[0104] An image acquisition module 301, configured to acquire a training image set; the training image set includes multiple training images, and each training image includes first annotation information and / or second annotation information. The first annotation information includes a first detection frame corresponding to a detection target in the training image, and the second annotation information includes a second detection frame corresponding to an obstacle in the training image;

[0105] A processing module 302, configured to input the training images in the training image set into the detection model for processing, and obtain a third detection frame and a first confidence level corresponding to the detection target in the training image output by the detection model, and / or a fourth detection frame and a second confidence level corresponding to the obstacle in the training image;

[0106] An adjustment module 303, configured to, when it is determined that the first confidence level corresponding to the detection target is greater than a first preset threshold, and / or the second confidence level corresponding to the obstacle is greater than a second preset threshold, adjust the model parameters of the detection model according to the first detection frame and the third detection frame corresponding to the detection target, and / or the second detection frame and the fourth detection frame corresponding to the obstacle, and return to the execution of the processing module 302 until the detection model meets a preset condition, and obtain a trained detection model. The first preset threshold is different from the second preset threshold.

[0107] In one embodiment, the adjustment module 303 is specifically configured to determine the detection accuracy corresponding to the detection model according to the first detection frame and the third detection frame corresponding to the detection target, and / or the second detection frame and the fourth detection frame corresponding to the obstacle; when it is determined that the detection accuracy is less than or equal to a third preset threshold, adjust the model parameters of the detection model.

[0108] In some embodiments, the first detection frame is a detection frame corresponding to the detection target that is completely displayed in the training image; the second detection frame is a detection frame corresponding to the obstacle that is completely displayed or partially displayed in the training image.

[0109] In some embodiments, the first preset threshold is greater than the second preset threshold.

[0110] Corresponding to the detection model training method described in the above embodiments, Figure 4The block diagram of the object detection device provided by the embodiment of the present application is shown. For the sake of convenience of description, only the parts related to the embodiment of the present application are shown.

[0111] Referring to Figure 4 , the device may include:

[0112] An image acquisition module 401, configured to acquire a first image corresponding to a target area;

[0113] An object detection module 402, configured to input the first image into a trained detection model for processing, and obtain a detection result output by the detection model;

[0114] Wherein, the detection model is trained based on the training method of the foregoing detection model.

[0115] In some embodiments, the detection result includes the category corresponding to one or more detected objects, the detection frame corresponding to each detected object, and the confidence corresponding to each detected object.

[0116] In one embodiment, the device further includes:

[0117] A first retention module, configured to obtain a first confidence corresponding to the detection target when it is determined that the detected object includes a detection target according to the category corresponding to the detected object; and retain the detection frame corresponding to the detection target when it is determined that the first confidence is greater than a first preset threshold.

[0118] In another embodiment, the device further includes:

[0119] A second retention module, configured to obtain a second confidence corresponding to the obstacle when it is determined that the detected object includes an obstacle according to the category corresponding to the detected object; and retain the detection frame corresponding to the obstacle when it is determined that the second confidence is greater than a second preset threshold.

[0120] Exemplarily, the first preset threshold is greater than the second preset threshold.

[0121] It should be noted that for the information interaction, execution process, etc. between the above-mentioned device / units, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, please refer to the method embodiment part for details, and will not be elaborated here.

[0122] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0123] Figure 5 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. As Figure 5 shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 only one is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50. When the processor 50 executes the computer program 52, the steps in the foregoing method embodiments of any of the above detection model training methods are implemented. Alternatively, when the processor 50 executes the computer program 52, the steps in the foregoing method embodiments of any of the above object detection methods are implemented.

[0124] The electronic device 5 may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art can understand that Figure 5 merely an example of the electronic device 5, which does not constitute a limitation on the electronic device 5. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0125] The processor 50 may be a central processing unit (CPU), and the processor 50 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0126] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as the hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk equipped on the electronic device 5, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 51 may also include both the internal storage unit and the external storage device of the electronic device 5. The memory 51 is used to store an operating system, application programs, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program, etc. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0127] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by an electronic device, the electronic device is caused to implement the steps in the embodiments of the above-mentioned various training methods of the detection model.

[0128] The embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. When the computer program is executed by an electronic device, the electronic device is caused to implement the steps in the embodiments of the above-mentioned various object detection methods.

[0129] The embodiment of the present application provides a computer program product. The computer program product includes a computer program. When the computer program is executed by an electronic device, the electronic device is caused to implement the steps in the embodiments of the above-mentioned various training methods of the detection model.

[0130] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by an electronic device, the electronic device is enabled to implement the steps in the above-mentioned embodiments of each object detection method.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, a computer program can be used to instruct relevant hardware to complete. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can at least include: any entity or device that can carry the computer program code to the device / electronic device, recording medium, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc. In some jurisdictions, according to legislation and patent practice, the computer-readable storage medium cannot be an electrical carrier signal and a telecommunication signal.

[0132] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0133] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0134] In the embodiments provided in the present application, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical, or other form.

[0135] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0136] The above-described embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application and should all be included within the protection scope of the present application.

Claims

1. A training method for a detection model, characterized in that, The method includes: Obtaining a training image set; the training image set includes multiple training images, and each training image includes first annotation information and / or second annotation information. The first annotation information includes a first detection box corresponding to a detection target in the training image, and the second annotation information includes a second detection box corresponding to an obstacle in the training image; Inputting the training images in the training image set into a detection model for processing to obtain a third detection box and a first confidence level corresponding to the detection target in the training image output by the detection model, and / or a fourth detection box and a second confidence level corresponding to the obstacle in the training image; When it is determined that the first confidence level corresponding to the detection target is greater than a first preset threshold, and / or it is determined that the second confidence level corresponding to the obstacle is greater than a second preset threshold, adjusting the model parameters of the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle, and returning to execute the step of inputting the training images in the training image set into the detection model for processing to obtain the third detection box and the first confidence level corresponding to the detection target in the training image output by the detection model, and / or the fourth detection box and the second confidence level corresponding to the obstacle in the training image, and subsequent steps until the detection model meets the preset conditions to obtain a trained detection model, where the first preset threshold is different from the second preset threshold.

2. The method according to claim 1, characterized in that, The adjusting the model parameters of the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle includes: Determining the detection accuracy corresponding to the detection model according to the first detection box and the third detection box corresponding to the detection target, and / or the second detection box and the fourth detection box corresponding to the obstacle; When it is determined that the detection accuracy is less than or equal to a third preset threshold, adjusting the model parameters of the detection model.

3. The method according to claim 1, characterized in that The first detection box is a detection box corresponding to the detection target that is completely displayed in the training image; the second detection box is a detection box corresponding to the obstacle that is completely displayed or partially displayed in the training image.

4. The method according to any one of claims 1 to 3, characterized in that, The first preset threshold is greater than the second preset threshold.

5. An object detection method, characterized in that, The method includes: Obtaining a first image corresponding to a target area; Inputting the first image into the trained detection model for processing to obtain a detection result output by the detection model; Wherein, the detection model is trained based on the training method of the detection model according to any one of claims 1-4.

6. The method according to claim 5, characterized in that The detection result includes the category of one or more detected objects, the detection box corresponding to each detected object, and the confidence level corresponding to each detected object; After the step of inputting the first image into the trained detection model for processing to obtain the detection result output by the detection model, the method further includes: When it is determined that the detected object includes a detection target according to the category of the detected object, obtain the first confidence corresponding to the detection target; When it is determined that the first confidence is greater than the first preset threshold, retain the detection box corresponding to the detection target.

7. The method according to claim 6, characterized in that, After inputting the first image into the trained detection model for processing to obtain the detection result output by the detection model, the method further includes: When it is determined that the detected object includes an obstacle according to the category of the detected object, obtain the second confidence corresponding to the obstacle; When it is determined that the second confidence is greater than the second preset threshold, retain the detection box corresponding to the obstacle.

8. The method according to claim 7, wherein The first preset threshold is greater than the second preset threshold.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the electronic device implements the training method of the detection model according to any one of claims 1 to 4, or the electronic device implements the object detection method according to any one of claims 5 to 8.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by an electronic device, the electronic device implements the training method of the detection model according to any one of claims 1 to 4, or the electronic device implements the object detection method according to any one of claims 5 to 8.