Visual inspection model training method and device, electronic equipment, computer readable storage medium and computer program product

By calculating the uncertainty and diversity scores of the image samples collected by the sweeping robot, high-quality samples were selected for model training, which solved the problems of high labeling costs and overfitting of the model in the existing technology, and improved the performance and generalization capabilities of the visual detection model.

CN120219807APending Publication Date: 2025-06-27UBTECH ROBOTICS CORP LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510225107.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

During the training process, the existing visual detection model of sweeping robots has high labeling costs and overfitting the model, and poor generalization ability.

Method used

By performing category detection and feature extraction on image samples collected by sweeping robots, uncertainty scores and diversity scores are calculated and fused into comprehensive scores, so that high-quality image samples can be screened for model training.

Benefits of technology

This method can effectively reduce labeling costs, eliminate duplicate samples, and improve the performance and generalization capabilities of the visual detection model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120219807A_ABST
    Figure CN120219807A_ABST
Patent Text Reader

Abstract

The invention provides a visual inspection model training method and device, electronic equipment, a computer program product and a computer readable storage medium. The method comprises the following steps: performing category detection on a plurality of image samples collected by the sweeping robot through a visual detection model of the sweeping robot to obtain an uncertainty score of each image sample; performing feature extraction on each image sample through a visual detection model to obtain image features of each image sample, and determining a diversity score of each image sample based on a first similarity between the image features of each image sample; fusing the uncertainty score and the diversity score of each image sample to obtain a comprehensive score of each image sample; and screening out a first image sample from the plurality of image samples based on the comprehensive score, and training a visual detection model based on the first image sample. According to the invention, the performance of the visual detection model obtained through training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to machine learning technology, and in particular, to a method, device, electronic device, computer program product, and computer-readable storage medium for training a visual detection model. Background Art

[0002] With the rapid development of artificial intelligence technology, floor-sweeping robots have gradually become an important part of smart homes. When a floor-sweeping robot starts to work, it often needs to first collect images of the working environment, and then train the visual detection model in the floor-sweeping robot based on the collected images, so that the floor-sweeping robot can identify obstacles in the working environment. However, since the images collected by the floor-sweeping robot often include multiple images in the same scene, the number of training samples is too large, resulting in too high a labeling cost for the training samples. Moreover, since the training samples include a large amount of repetitive content, the trained visual detection model is overfitted and has poor generalization ability. Summary of the Invention

[0003] Embodiments of the present application provide a method, device, electronic device, computer program product, and computer-readable storage medium for training a visual detection model, which can improve the performance of the trained visual detection model.

[0004] The technical solution of the embodiments of the present application is implemented as follows:

[0005] Embodiments of the present application provide a method for training a visual detection model, the method including:

[0006] Performing class detection on a plurality of image samples collected by the floor-sweeping robot through the visual detection model of the floor-sweeping robot to obtain the uncertainty scores of the respective image samples;

[0007] Performing feature extraction on the respective image samples through the visual detection model to obtain the image features of the respective image samples, and determining the diversity scores of the respective image samples based on the first similarity between the image features of the respective image samples;

[0008] Fusing the uncertainty scores and the diversity scores of the respective image samples to obtain the comprehensive scores of the respective image samples;

[0009] Screening out first image samples from the plurality of image samples based on the comprehensive scores, and training the visual detection model based on the first image samples.

[0010] Embodiments of the present application provide a training device for a visual detection model, including:

[0011] A detection module, configured to perform class detection on a plurality of image samples collected by the floor cleaning robot through a visual detection model of the floor cleaning robot, so as to obtain the uncertainty scores of the image samples;

[0012] A determination module, configured to extract features of each of the image samples through the visual detection model to obtain the image features of each of the image samples, and determine the diversity scores of the image samples based on the first similarity between the image features of the image samples;

[0013] A fusion module, configured to fuse the uncertainty scores and diversity scores of the image samples to obtain the comprehensive scores of the image samples;

[0014] A training module, configured to screen out first image samples from the plurality of image samples based on the comprehensive scores, and train the visual detection model based on the first image samples.

[0015] In some embodiments, the above detection module is further configured to perform class detection on each of the image samples collected by the floor cleaning robot to obtain a plurality of detection frames of the image sample, and one detection frame indicates the class to which a target object in the image sample belongs; determine the class uncertainty and confidence uncertainty of each detection frame; determine the uncertainty score of each detection frame based on the class uncertainty and confidence uncertainty of each detection frame; and average the uncertainty scores of each detection frame to obtain the uncertainty score of the image sample.

[0016] In some embodiments, the above detection module is further configured to, for each of the detection frames, determine the total number of classes of the target objects included in the image sample and the probability that the target object in the detection frame is predicted to be each of the classes based on the class to which the target object indicated by the detection frame belongs; multiply the probability that the target object in the detection frame is predicted to be each of the classes by the logarithm of the corresponding probability to obtain the uncertainty that the target object in the detection frame is predicted to be each of the classes; and sum the uncertainties that the target object in the detection frame is predicted to be each of the classes to obtain the class uncertainty of the detection frame.

[0017] In some embodiments, the above detection module is further configured to, for each of the detection frames, determine a first weight of the class uncertainty and a second weight of the confidence uncertainty; fuse the class uncertainty and the confidence uncertainty based on the first weight, the second weight, and the total number of classes to obtain the uncertainty score of the detection frame; wherein the uncertainty score is positively correlated with the class uncertainty and negatively correlated with the total number of classes and the confidence uncertainty.

[0018] In some embodiments, the above-mentioned determination module is further configured to, for each of the image samples, determine a first similarity between the image features of the image sample and the image features of other images; determine the difference between a preset value and each of the first similarities, and sum up the differences to obtain the diversity score of the image sample.

[0019] In some embodiments, the above-mentioned training module is further configured to sort the multiple image samples in ascending order according to the comprehensive score to obtain an image sample sequence; screen out a first number of image samples from the front end of the image sample sequence, and determine the screened first number of image samples as the first image samples.

[0020] In some embodiments, the above-mentioned training module is further configured to, based on the comprehensive score, screen out a second number of image samples whose comprehensive score exceeds a score threshold from the multiple image samples, and determine the screened second number of image samples as the first image samples.

[0021] In some embodiments, the above-mentioned training module is further configured to, in response to the number of the first image samples being multiple, determine a second similarity between the first image samples; screen out image sample pairs corresponding to the second similarity that exceed a similarity threshold from the multiple first image samples; perform a deduplication process on the screened image sample pairs to obtain second image samples; and train the visual detection model based on the second image samples and third image samples other than the image sample pairs in the first image samples.

[0022] In some embodiments, the above-mentioned training module is further configured to use the new visual detection model obtained by training the visual detection model based on the first image samples to perform class detection on fourth image samples other than the first image samples in the multiple image samples to obtain the uncertainty scores of the fourth image samples; use the new visual detection model to extract the image features of the fourth image samples to obtain the image features of the fourth image samples, and determine the diversity scores of the fourth image samples based on the third similarity between the image features of the fourth image samples; fuse the uncertainty scores and diversity scores of the fourth image samples to obtain the comprehensive scores of the fourth image samples; screen out fifth image samples from the fourth image samples based on the comprehensive scores, and perform iterative training on the new visual detection model based on the fifth image samples until the total number of image samples participating in the training reaches a quantity threshold and then stop the training.

[0023] An embodiment of the present application provides an electronic device, and the electronic device includes:

[0024] A memory for storing computer-executable instructions or computer programs;

[0025] A processor, when executing the computer-executable instructions or computer programs stored in the memory, implements the training method of the visual detection model provided by the embodiments of the present application.

[0026] The embodiments of the present application provide a computer-readable storage medium storing a computer program or computer-executable instructions, which when executed by a processor, implement the training method of the visual detection model provided by the embodiments of the present application.

[0027] The embodiments of the present application provide a computer program product including a computer program or computer-executable instructions, which when executed by a processor, implement the training method of the visual detection model provided by the embodiments of the present application.

[0028] The embodiments of the present application have the following beneficial effects:

[0029] Through the visual detection model of the sweeping robot, category detection is performed on multiple image samples collected by the sweeping robot to obtain the uncertainty scores of each image sample. By obtaining the uncertainty scores of the image samples through the visual detection model of the sweeping robot, a score for measuring the difficulty of category detection of the image samples can be obtained, thereby improving the accuracy of the comprehensive score obtained based on the uncertainty scores subsequently; through the visual detection model, feature extraction is performed on each image sample to obtain the image features of each image sample, and based on the first similarity between the image features of each image sample, the diversity scores of each image sample are determined. By obtaining the diversity scores of the image samples through the visual detection model of the sweeping robot, a score for measuring whether the image samples are duplicate samples can be obtained, thereby improving the accuracy of the comprehensive score obtained based on the diversity scores subsequently. Then, the uncertainty scores and diversity scores of each image sample are fused to obtain the comprehensive scores of each image sample. Based on the comprehensive scores, the first image samples are selected from multiple image samples, and the visual detection model is trained based on the first image samples. By using the comprehensive scores to represent the quality of the image samples, duplicate samples and high-quality first image samples that are easily recognizable can be eliminated, thereby improving the performance of the visual detection model trained based on the first image samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 is a schematic architecture diagram of the training system 100 of the visual detection model provided by the embodiments of the present application;

[0031] Figure 2 is a schematic structural diagram of the electronic device 500 provided by the embodiments of the present application;

[0032] Figure 3A is the first process schematic diagram of the training method of the visual detection model provided by the embodiments of the present application;

[0033] Figure 3B is the second process schematic diagram of the training method of the visual detection model provided by the embodiments of the present application;

[0034] Figure 3C is the third process schematic diagram of the training method of the visual detection model provided by the embodiments of the present application.

[0035] It should be noted that the above "first" and "second" are only used to distinguish different solutions, and do not represent the distinction of the advantages and disadvantages of the solutions or the priority in the implementation process. Detailed implementation manners

[0036] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0037] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0038] In the following description, the terms "first / second / third" are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0039] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.

[0040] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meaning as commonly understood by those of ordinary skill in the art to which the present application belongs. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0041] In the embodiments of this application, when collecting and processing relevant data in practical applications, it should strictly comply with the requirements of relevant laws and regulations, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope authorized by laws and regulations and the personal information subject.

[0042] Before further elaborating on the embodiments of this application, the nouns and terms involved in the embodiments of this application are explained. The nouns and terms involved in the embodiments of this application are applicable to the following explanations.

[0043] 1) Floor cleaning robot: A floor cleaning robot is usually equipped with sensors such as collision sensors, infrared sensors, or laser navigation systems to sense the surrounding environment and obstacles. The floor cleaning robot can also plan a cleaning path, avoid furniture and other obstacles. The floor cleaning robot uses navigation technologies (such as laser navigation, visual navigation, etc.) to build an indoor map and efficiently cover the cleaning area. The floor cleaning robot can identify and avoid obstacles to prevent falling from a height.

[0044] 2) Visual detection model: The visual detection model is embedded in the floor cleaning robot and uses computer vision technology to identify, locate, and classify target objects in the images or videos collected by the floor cleaning robot.

[0045] With the rapid development of artificial intelligence technology, floor cleaning robots have gradually become an important part of smart homes. Their core functions rely on visual perception capabilities, including obstacle detection, furniture recognition, and floor material judgment. Through artificial intelligence visual technology, the floor cleaning robot can perform path planning, obstacle avoidance, and scene recognition more efficiently, thereby improving the cleaning efficiency and user experience.

[0046] However, in the actual application of floor cleaning robots, the diversity and complexity of scenes pose higher requirements for the visual detection model. For example, in different indoor environments, the types of obstacles are diverse and change frequently, and there may also be significant differences in lighting conditions. Therefore, in order to maintain the accuracy and robustness of the visual detection model of the floor cleaning robot, it is usually necessary to collect data from new scenes and annotate them for training. However, there are the following main problems in the data return process:

[0047] (1) High data annotation cost: The images collected by the floor cleaning robot usually contain a large number of categories that need to be annotated, including floor materials, furniture, and various obstacles. Thoroughly annotating these data requires a large amount of manpower and time.

[0048] (2) Serious data redundancy problem: During the operation of the floor cleaning robot, consecutive frame images will be collected, or the same area will be passed through multiple times, resulting in a large amount of similar or duplicate information in the returned data.

[0049] To solve the above problems, there is an urgent need for an effective data screening method to select the most valuable samples from the feedback data, which can not only significantly reduce the annotation cost but also maximize the performance of the visual detection model.

[0050] The embodiments of the present application provide a training method, device, electronic device, computer-readable storage medium, and computer program product for a visual detection model, which can improve the performance of the trained visual detection model. The following describes the exemplary application of the training device for the visual detection model provided by the embodiments of the present application. The device provided by the embodiments of the present application can be implemented as various types of terminals such as laptop computers, tablet computers, desktop computers, set-top boxes, smartphones, smart speakers, smart watches, smart TVs, vehicle terminals, floor-sweeping robots, etc., or can be implemented as a server. The following will describe the exemplary application when the device is implemented as a server.

[0051] See Figure 1 , Figure 1 is a schematic diagram of the architecture of the training system 100 for the visual detection model provided by the embodiments of the present application. To support the training application of a visual detection model, the terminal 400 is connected to the server 200 through the network 300. The network 300 can be a wide area network, a local area network, or a combination of the two.

[0052] The terminal 400 (such as a floor-sweeping robot) is used to collect a plurality of image samples and transmit the image samples to the server 200 through the network 300.

[0053] The server 200 is used to perform class detection on a plurality of image samples collected by the floor-sweeping robot through the visual detection model of the floor-sweeping robot to obtain the uncertainty scores of each image sample, extract the image features of each image sample through the visual detection model to obtain the image features of each image sample, determine the diversity scores of each image sample based on the first similarity between the image features of each image sample, fuse the uncertainty scores and diversity scores of each image sample to obtain the comprehensive scores of each image sample, screen out the first image samples from the plurality of image samples based on the comprehensive scores, and train the visual detection model based on the first image samples.

[0054] In some embodiments, the server 200 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The terminal and the server may be directly or indirectly connected through wired or wireless communication methods, which are not limited in the embodiments of the present application.

[0055] See Figure 2 , Figure 2 is a schematic structural diagram of the electronic device 500 provided by the embodiments of the present application. Taking the first server 200 in Figure 1 as an example, Figure 2 the electronic device 500 shown includes: at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. Each component in the electronic device 500 is coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 2 all kinds of buses are labeled as the bus system 540.

[0056] The processor 510 may be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a Digital Signal Processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or any conventional processor, etc.

[0057] The user interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as keyboards, mice, microphones, touch screen displays, cameras, and other input buttons and controls.

[0058] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memories, hard disk drives, optical disc drives, etc. The memory 550 optionally includes one or more storage devices that are physically located away from the processor 510.

[0059] The memory 550 includes volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.

[0060] In some embodiments, the memory 550 is capable of storing data to support various operations. Examples of such data include programs, modules, and data structures, or subsets or supersets thereof, which are exemplarily described below.

[0061] The operating system 551 includes system programs for processing various basic system services and performing hardware-related tasks, such as the framework layer, the core library layer, the driver layer, etc., for implementing various basic services and processing hardware-based tasks;

[0062] The network communication module 552 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 520. Exemplary network interfaces 520 include: Bluetooth, wireless fidelity (WiFi), and universal serial bus (USB), etc.;

[0063] The presentation module 553 is used to enable the presentation of information (such as a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 associated with the user interface 530 (such as a display screen, a speaker, etc.);

[0064] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one of one or more input devices 532.

[0065] In some embodiments, the device provided by the embodiments of the present application may be implemented in software. Figure 2 Shown is a training device 555 for a visual detection model stored in the memory 550, which may be software in the form of programs and plugins, etc., including the following software modules: a detection module 5551, a determination module 5552, a fusion module 5553, and a training module 5554. These modules are logical, and thus can be arbitrarily combined or further split according to the functions to be implemented. The functions of each module will be described below.

[0066] In some other embodiments, the device provided by the embodiments of the present application can be implemented in a hardware manner. As an example, the device provided by the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the training method of the visual detection model provided by the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs) or other electronic components.

[0067] Next, the training method of the visual detection model provided by the embodiments of the present application will be described. As before, the electronic device for implementing the training method of the visual detection model provided by the embodiments of the present application can be a terminal, a server, or a combination of both. Therefore, the execution subject of each step will not be repeated hereinafter.

[0068] See Figure 3A , Figure 3A is the first flowchart of the training method of the visual detection model provided by the embodiments of the present application, and will be described in conjunction with the steps shown in Figure 3A shown.

[0069] In step 101, through the visual detection model of the sweeping robot, category detection is performed on a plurality of image samples collected by the sweeping robot, and the uncertainty score of each image sample is obtained.

[0070] In actual implementation, the sweeping robot is usually equipped with sensors, such as collision sensors, infrared sensors or laser navigation systems, to sense the surrounding environment and obstacles. The sweeping robot can also plan a cleaning path to avoid furniture and other obstacles. The sweeping robot uses navigation technologies (such as laser navigation, visual navigation, etc.) to build an indoor map and efficiently cover the cleaning area. The sweeping robot can identify and avoid obstacles to prevent falling from a height.

[0071] In actual implementation, the visual detection model is a visual detection model deployed in the sweeping robot. The visual detection model is a visual detection model that uses computer vision technology to identify, locate and classify target objects in images or videos. The visual detection model is widely used in industrial automation, security monitoring, medical diagnosis, driverless and other fields.

[0072] In some embodiments, the class detection of multiple image samples collected by the floor cleaning robot in step 101 to obtain the uncertainty scores of each image sample can be implemented through steps 1011 to 1014 as follows: Figure 3B shown in:

[0073] In step 1011, for each image sample collected by the floor cleaning robot, perform class detection on the image sample to obtain multiple detection boxes of the image sample. One detection box indicates the class to which a target object in the image sample belongs.

[0074] In actual implementation, the way for the floor cleaning robot to collect each image sample is usually to collect images of the surrounding environment of the path passed by the floor cleaning robot through an image acquisition device (such as a camera) deployed in the floor cleaning robot, and use the collected images of the surrounding environment as image samples.

[0075] In actual implementation, the image sample can be input into a visual detection model. The feature extraction network in the visual detection model extracts the features of the image sample to obtain the image features of the image sample, and then inputs the image features into a fully connected layer. The fully connected layer outputs the class of the target object included in the image sample and the detection box surrounding the target object.

[0076] As an example, assume that there are three target objects in the image sample, such as a table, a chair, and a toy car. Then, the detection results obtained by performing class detection on the image sample include three detection boxes, namely the detection box of the table, the detection box of the chair, and the detection box of the toy car. At the same time, after identifying the detection boxes of the target objects included in the image sample, mark the class to which the target object in the detection box belongs. Continuing with the above example, for example, mark class 1 on the detection box of the table, where class 1 indicates that the target object in this detection box is a table. Similarly, mark class 2 indicating that the target object in this detection box is a chair on the detection box of the chair, and mark class 3 indicating that the target object in this detection box is a toy car on the detection box of the toy car.

[0077] Through image class detection, the floor cleaning robot can better perceive and understand its surrounding environment, identify different types of target objects, thereby improving its environmental adaptability, helping the floor cleaning robot to perform path planning more accurately, avoid collisions and obstacles, and improve the cleaning efficiency.

[0078] In step 1012, determine the class uncertainty and confidence uncertainty of each detection box.

[0079] In actual implementation, the confidence of a detection box can be an indicator representing the degree of trust in the detection result. The confidence of each detection box can be composed of two parts: the confidence of the existence of an object and the confidence of the object category. The confidence of the existence of an object reflects the credibility of the visual detection model in judging that there is indeed an object in the bounding box, while the confidence of the object category is the credibility of the visual detection model in determining that the object within the bounding box belongs to a specific category. The product of the two is the total confidence of the detection box.

[0080] As an example, a way to determine the confidence uncertainty of a detection box can be that while the visual detection model predicts the detection boxes of each target object, it also outputs the confidence of each detection box, and directly uses the output confidence as the confidence uncertainty.

[0081] In actual implementation, first, the image features of the image sample are extracted through the feature extraction network of the visual detection model. These image features include, but are not limited to, information such as the shape, texture, and position of the objects (i.e., target objects) in the image sample. Then, the visual detection model will first generate a series of candidate regions, which may be the potential positions of the objects. For each candidate region, the visual detection model will predict the category of the object in the candidate region, such as predicting the scores of one or more categories for each position on each candidate region. These scores represent the confidence of the visual detection model in the appearance of each category. The visual detection model will also predict the coordinates of the bounding box, which define the position and size of the object detected in the image sample. Among them, the confidence is usually determined by combining the category scores and the results of bounding box regression. Specifically, the confidence can be obtained by multiplying the category scores and the accuracy of the bounding box.

[0082] In some embodiments, the determination of the category uncertainty of each detection box in step 1012 can be achieved through the following technical solution: for each detection box, based on the category to which the detection box indicates the target object belongs, determine the total number of categories of the target objects included in the image sample, and the probabilities that the target object in the detection box is predicted to be each category; multiply the probabilities that the target object in the detection box is predicted to be each category and the logarithms of the corresponding probabilities to obtain the uncertainties that the target object in the detection box is predicted to be each category; sum up the uncertainties that the target object in the detection box is predicted to be each category to obtain the category uncertainty of the detection box.

[0083] In actual implementation, if the probabilities corresponding to the categories to which the objects indicated by the detection frames belong are relatively close, it indicates that it is difficult to determine the specific category of the objects in the detection frames, and the category uncertainty of the detection frames is relatively low. For example, the probability that the object in detection frame A belongs to the category of chair is 0.25, the probability that the object in detection frame A belongs to the category of table is 0.25, the probability that the object in detection frame A belongs to the category of sofa is 0.25, and the probability that the object in detection frame A belongs to the category of clothes hanger is 0.25. At this time, it is difficult to determine which category the object belongs to.

[0084] In actual implementation, to determine the category uncertainty of the detection frame, specifically, the following formula (1) can be referred to:

[0085]

[0086] In formula (1), C is the total number of categories of the object, P(y cls |b i ) is the probability that the object predicted by detection frame b i belongs to category cls of the object, b i is the i-th detection frame, and H(b i ) is the category uncertainty of detection frame b i .

[0087] According to formula (1), it can be known that when calculating the category uncertainty of each detection frame, the probabilities that the object in the detection frame is predicted to belong to various categories and their logarithms are multiplied to obtain the information entropy of the detection frame as the uncertainty that the object in the corresponding detection frame is predicted to belong to various categories. Information entropy is an index to measure the uncertainty of probability distribution. Generally, the higher the probability, the smaller its logarithm, and the smaller the result after multiplication, indicating the lower the uncertainty; on the contrary, the lower the probability, the larger its logarithm, and the larger the result after multiplication, indicating the higher the uncertainty.

[0088] Then, the uncertainties that the object in the detection frame is predicted to belong to various categories are added to obtain the category uncertainty of the detection frame. The higher this value is, the more uncertain the visual detection model is about the category prediction of the object; on the contrary, the lower the value is, the more certain the visual detection model is about the category prediction of the object.

[0089] Through the above method, it is possible to combine the probabilities that the object in the detection frame belongs to various categories to accurately determine the uncertainty score of the detection frame, improving the accuracy of the determined uncertainty score.

[0090] In step 1013, based on the category uncertainty and confidence uncertainty of each detection frame, the uncertainty score of each detection frame is determined.

[0091] In some embodiments, determining the uncertainty score of each detection box based on the category uncertainty and confidence uncertainty of each detection box in step 1013 can be achieved through steps 10131 to 10132 as shown in Figure 3C :

[0092] In step 10131, for each detection box, determine the first weight of the category uncertainty and the second weight of the confidence uncertainty.

[0093] In actual implementation, the first weight and the second weight can be values set according to experience. The first weight and the second weight corresponding to each detection box can be the same, or different first weights and second weights can be set according to the different objects included in the detection box. For example, if the object included in the inspection box is a chair, the first weight can be 0.4 and the second weight can be 0.6. If the object included in the inspection box is a flying insect, the first weight can be 0.7 and the second weight can be 0.3.

[0094] In the above manner, by setting the first weight and the second weight, the degree of attention to the category uncertainty and the confidence uncertainty can be adjusted, and an image sample more suitable for training the visual detection model can be obtained.

[0095] In step 10132, based on the first weight, the second weight, and the total number of categories, fuse the category uncertainty and the confidence uncertainty to obtain the uncertainty score of the detection box; wherein, the uncertainty score is positively correlated with the category uncertainty and negatively correlated with the total number of categories and the confidence uncertainty.

[0096] In actual implementation, high-quality image samples are often image samples with high confidence of the detection boxes included and concentrated prediction probabilities of the categories corresponding to the detection boxes. For example, if the confidence of the detection box included in image sample A is 0.9 and the probability that the category predicted by this detection box is a chair is 0.95, then image sample A is a high-quality image sample. At the same time, since the subsequent sorting of the image samples is in ascending order and the image samples are selected from the front end of the image sample sequence for training the visual detection model, the determined uncertainty score of the image sample is positively correlated with the uncertainty and negatively correlated with the confidence uncertainty.

[0097] In actual implementation, the method for determining the uncertainty score of the detection box can refer to the following formula (2):

[0098]

[0099] In formula (2), U(b i ) is the uncertainty score of detection box b i , α is the first weight, H(b i ) is detection box bi The category uncertainty, β is the second weight, C is the total number of categories of the target object, s i is the confidence of the detection box b i in the confidence.

[0100] In the above way, by identifying and adjusting the detection boxes with higher uncertainty, the results of false positives (false alarms) and false negatives (missed detections) can be reduced, thereby improving the overall detection accuracy. The uncertainty score provides an explanation for the uncertainty of the prediction results of the visual detection model, which helps users understand the decision-making process of the visual detection model. Different application scenarios have different tolerances for uncertainty. The method of fusing uncertainty can adjust the weights according to the requirements of specific scenarios to better adapt to various applications.

[0101] In step 1014, the uncertainty scores of each detection box are averaged to obtain the uncertainty score of the image sample.

[0102] In actual implementation, after determining the uncertainty scores of each detection box included in the image sample, the uncertainty scores of each detection box can be averaged, and then the obtained average value can be used as the uncertainty score of the image sample. Similarly, the uncertainty scores of the detection boxes can also be weighted and summed according to the size of the target object included in the detection box.

[0103] In actual implementation, the method for determining the uncertainty score of the image sample can refer to the following formula (3):

[0104]

[0105] In formula (3), U(X) is the uncertainty score of the image sample X, n is the number of detection boxes, U(b i ) is the uncertainty score of the detection box b i in the confidence.

[0106] In the above way, it is possible to fuse whether the image sample is easy to be recognized and the confidence of the recognized detection box, obtain an accurate uncertainty score that can characterize whether the quality of the image sample itself is sufficient, and thereby improve the quality of the subsequent selected image samples.

[0107] In step 102, through the visual detection model, feature extraction is performed on each image sample to obtain the image features of each image sample, and based on the first similarity between the image features of each image sample, the diversity score of each image sample is determined.

[0108] In actual implementation, an image sample can be input into a visual detection model, and the feature extraction network of the visual detection model extracts features from the image sample to obtain the image features of the image sample. Among them, the feature extraction network abstracts the image through convolutional layers and pooling layers, and learns features that can represent high-level attributes such as object shape and texture from the original pixel data. The feature extraction network is designed with a multi-layer structure, and each layer of convolutional operation can capture features at different scales in the image. Generally, the shallow network captures low-level features such as edges and textures, while the deep network can capture more abstract high-level features, such as the overall shape or part of an object. As the number of network layers increases, the dimension of the feature map usually gradually decreases, while the number of channels may increase. Such a design helps to reduce the computational complexity while retaining key information.

[0109] In some embodiments, determining the diversity score of each image sample based on the first similarity between the image features of each image sample can be achieved through the following technical solution: for each image sample, determine the first similarity between the image features of the image sample and the image features of other images; determine the difference between the preset value and each first similarity, and sum up each difference to obtain the diversity score of the image sample.

[0110] In actual implementation, the way to determine the first similarity between the image features of an image sample and the image features of other images can be to determine the cosine similarity, Euclidean distance, etc. between the image features of the image sample and the image features of other images. It should be noted that the way to determine the first similarity can be to directly determine the similarity between the image features of any one image sample and the image features of all other image samples. For example, if the image samples include image sample A, image sample B, and image sample C, then the first similarity D between image sample A and image sample B and the first similarity E between image sample A and image sample C can be determined. At the same time, the way to determine the first similarity can also be to determine the similarity between any one image sample and the image sample whose shooting time interval from this image sample is greater than the preset time interval threshold. For example, the image samples include image sample A, image sample B, and image sample C, where the shooting time of image sample A is 12:15, the shooting time of image sample B is 12:17, the shooting time of image sample C is 12:19, and the time interval threshold is 3 minutes, then the first similarity between image sample A and image sample B can be determined.

[0111] In actual implementation, the first similarity between the image features of an image sample and the image features of other images can be referred to the following formula (4):

[0112]

[0113] In formula (4), Si,j For image sample x i and image sample x j the similarity between them, f(x i ) is the image feature of image sample x i and f(x j ) is the image feature of image sample x j .

[0114] In actual implementation, the diversity score of the image sample can be determined by referring to image formula (5):

[0115] D(x i ) = ∑ i≠j 1 - S i,j (5)

[0116] In formula (5), D(x i ) is the diversity score of image sample x i , S i,j is the similarity between image sample x i and image sample x j , and 1 is a preset value.

[0117] Through the above method, a multi - sample score can be obtained to measure whether the information included in the image samples is repeated, thereby improving the diversity of the subsequently selected image samples.

[0118] In step 103, the uncertainty scores and diversity scores of each image sample are fused to obtain the comprehensive scores of each image sample.

[0119] In actual implementation, after obtaining the uncertainty scores and diversity scores of the image samples, the comprehensive scores can be directly obtained by adding the uncertainty scores and diversity scores and using the added result as the comprehensive score, or by multiplying the uncertainty scores and diversity scores and using the multiplied result as the comprehensive score, or by performing weighted summation on the uncertainty scores and diversity scores and using the weighted summation result as the comprehensive score. Here is a specific implementation method:

[0120] The method for obtaining the comprehensive score can be seen in the following formula (6):

[0121] F(x) = U(X) + λD(X) (6)

[0122] In formula (6), F(x) is the comprehensive score of image sample x, D(x) is the diversity score of image sample x, λ is an adjustment parameter, and U(X) is the uncertainty score of image sample X.

[0123] In step 104, the first image sample is screened out from multiple image samples based on the comprehensive score, and the visual detection model is trained based on the first image sample.

[0124] In some embodiments, screening out the first image sample from multiple image samples based on the comprehensive score in step 104 can be achieved through the following technical solution: the multiple image samples are sorted in ascending order according to the comprehensive score to obtain an image sample sequence; the first number of image samples is screened out starting from the front end of the image sample sequence, and the screened first number of image samples is determined as the first image sample.

[0125] In actual implementation, the way to screen the first image sample can be to first sort the first image samples according to the comprehensive score of each image sample, and then obtain an image sample sequence. For example, if the comprehensive score of image sample A is 10, the comprehensive score of image sample B is 11, and the comprehensive score of image sample C is 9, the obtained image sample sequence is image sample B, image sample A, and image sample C.

[0126] In actual implementation, since the quality of the image samples at the front end of the image sample sequence is relatively high, in each training, a fixed number of image samples can be screened out from the front end of the image sample sequence. For example, if the image sample sequence includes 10 image samples and the first number is 5, the first 5 image samples in the image sample sequence can be used as the first image sample.

[0127] Through the above method, a fixed number of high-quality image samples can be screened out, thereby improving the quality of the obtained first image sample.

[0128] In some embodiments, screening out the first image sample from multiple image samples based on the comprehensive score in step 104 can also be achieved through the following technical solution: based on the comprehensive score, the second number of image samples with a comprehensive score exceeding the score threshold is screened out from multiple image samples, and the screened second number of image samples is determined as the first image sample.

[0129] In actual implementation, to ensure the quality of the screened image samples, the image samples in the image sample sequence with a comprehensive score exceeding the score threshold can be used as the first image sample. For example, if the image sample sequence is image sample B(11), image sample A(10), and image sample C(9), and the score threshold is 9.5, the second number can be determined to be 2, and the first image samples are image sample B and image sample A.

[0130] Through the above method, it can be ensured that the obtained first image samples are high-quality image samples.

[0131] In actual implementation, after obtaining high-quality image samples, the image samples can be labeled to obtain the categories of obstacles included in the image samples, and the image samples are input into the visual detection model. The visual detection model predicts the categories of obstacles included in the image samples. A loss function of the visual detection model is constructed based on the categories of obstacles predicted by the visual detection model and the label values of the image samples. Finally, the visual detection model is trained based on the loss function. Among them, the loss function can be a cosine loss function, a mean square error loss function, etc. The specific loss function can be set according to the actual situation.

[0132] In some embodiments, training the visual detection model based on the first image sample in step 104 can be implemented by the following technical solution: in response to the number of the first image samples being multiple, determining a second similarity between the first image samples; screening out the image sample pairs corresponding to the second similarities that exceed the similarity threshold from the multiple first image samples; performing a duplicate removal process on the screened image sample pairs to obtain second image samples; and training the visual detection model based on the second image samples and the third image samples in the first image samples except the image sample pairs.

[0133] In actual implementation, since the sweeping robot may capture multiple identical images when it is stationary, and in order to avoid using duplicate images to train the visual detection model, after obtaining the first image samples, the second similarity between the first image samples can be determined, and the image sample pairs with the second similarity exceeding the similarity threshold can be determined. Then, a duplicate removal process is performed on the image sample pairs to obtain second image samples. For example, if the image sample pair is the first image sample A and the second image sample B, at this time, the comprehensive scores of the first image sample A and the second image sample B can be determined. If the comprehensive score of the first image sample A exceeds the comprehensive score of the second image sample B, the second image sample B can be deleted, and the first image sample A is used as the second image sample. If the comprehensive scores of the first image sample A and the second image sample B are the same, any one can be deleted arbitrarily, and the undeleted first image sample is used as the second image sample.

[0134] In actual implementation, after obtaining the second image sample, the visual detection model can be trained based on the third image sample other than the image sample pair in the first image sample and the second image sample. For example, the first image sample includes the first image sample A, the first image sample B, the first image sample C, and the first image sample D. If the second similarity between the first image sample A and the first image sample B exceeds the similarity threshold, it is determined that the first image sample A and the first image sample B form an image sample pair. Then, the first image sample C and the first image sample D are the third image samples. If it is determined that the first image sample A is the second image sample, the visual detection model can be trained based on the first image sample A, the first image sample C, and the first image sample D.

[0135] Through the above method, duplicate samples in the training samples can be eliminated, avoiding the problem of overfitting of the trained visual detection model and improving the generalization of the visual detection model.

[0136] In some embodiments, after performing the training of the visual detection model based on the first image sample in step 104, the following technical solutions can also be executed: using the new visual detection model obtained by training the visual detection model based on the first image sample to perform class detection on the fourth image samples other than the first image sample in multiple image samples to obtain the uncertainty scores of each fourth image sample; using the new visual detection model to extract the image features of each fourth image sample to obtain the image features of each fourth image sample, and determining the diversity scores of each fourth image sample based on the third similarity between the image features of each fourth image sample; fusing the uncertainty scores and diversity scores of each fourth image sample to obtain the comprehensive scores of each fourth image sample; screening out the fifth image sample from the fourth image samples based on the comprehensive scores, and iteratively training the new visual detection model based on the fifth image sample until the total number of image samples participating in the training reaches the quantity threshold and then stopping the training.

[0137] In actual implementation, after training the visual detection model with the first image sample to obtain a new visual detection model, in order to further improve the performance of the visual detection model, the new visual detection model can be iteratively trained. Specifically, the training samples other than the first image sample that has already participated in the training can be input into the new visual detection model, and the comprehensive scores between each image sample are determined through the new visual detection model. Then, the fifth image sample for training the new visual detection model is screened based on the comprehensive scores, so as to train the new visual detection model based on the fifth image sample.

[0138] As an example, multiple image samples include image sample A, image sample B, and image sample C. If it is determined that the first image sample is image sample A during the first training, it indicates that image sample A has been used for training the visual detection model. Then, the unparticipated image samples B and C can be input into the new visual detection model to determine the comprehensive scores of image samples B and C through the new visual detection model. If it is determined through screening that the fifth image sample is image sample B, the new visual detection model can be trained based on image sample B.

[0139] In actual implementation, after obtaining a new visual detection model through each training, iterative training can be performed on the new visual detection model until the total number of image samples participating in the training reaches a quantity threshold. For example, the quantity threshold is 100, the number of image samples participating in the training during the first training is 10, the number of image samples participating in the training during the second training is 50, and the number of image samples participating in the training during the third training is 40. Then, the training of the visual detection model can be stopped after the third training.

[0140] Through the above iterative training method, the visual detection model is allowed to self-correct according to new data or errors in each iteration, thereby gradually improving the prediction accuracy rate, helping the visual detection model learn more generalized features, and enabling the visual detection model to perform better on new and unseen data.

[0141] Next, an exemplary application of the embodiments of the present application in an actual application scenario will be described.

[0142] Train a visual detection model for a floor sweeper based on existing labeled data to evaluate the sample scores of unlabeled data. For the unlabeled data set Data(u), the following steps are adopted to screen high learning value samples:

[0143] First, determine the uncertainty of each image sample. Assume that for an unlabeled sample x, the set of predicted detection frames of the visual detection model is B = {b1, b2......, b n}, where n is the number of detection frames. Then, the uncertainty score U(b i ) of the detection frame b i is calculated by the following formula:

[0144]

[0145] In formula (7), U(b i ) is the uncertainty score of the detection frame b i , α is the first weight, H(b i ) is the class uncertainty of the detection frame b i , β is the second weight, C is the total number of target object classes, and s iFor detection box b i 's confidence level.

[0146]

[0147] In formula (8), C is the total number of categories of the target object, and P(y cls |b i ) is the predicted probability of the predicted category cls of the target object by detection box b i , b i is the i-th detection box, and H(b i ) is the category uncertainty of detection box b i , representing the category uncertainty and used to measure the uncertainty of detection box bi in category prediction.

[0148] After determining the uncertainty score of each detection box included in the image sample x, the uncertainty score of the image sample x can be determined through the following formula:

[0149]

[0150] In formula (9), U(X) is the uncertainty score of the image sample X, n is the number of detection boxes, and U(b i ) is the uncertainty score of detection box b i .

[0151] After determining the uncertainty score of the image sample, the diversity of the image sample can be determined. Specifically, first, through the feature extraction network of the visual detection model, the feature vector (i.e., the above-mentioned image feature) f(x) of the sample x in the unlabeled dataset Du is extracted. The similarity matrix S (i.e., the above-mentioned first similarity) between image samples is defined as:

[0152]

[0153] In formula (10), S i,j is the similarity between image sample x i and image sample x j , f(x i ) is the image feature of image sample x i , and f(x j ) is the image feature of image sample x j . Among them, the larger S i,j is, the more similar image sample x i and image sample x j are.

[0154] Next, the diversity score D(x) of sample x is defined as:

[0155] D(x i ) = ∑i≠j 1-S i,j (11)

[0156] In formula (11), D(x i ) is the diversity score of the image sample x i , S i,j is the similarity between the image sample x i and the image sample x j , and 1 is a preset value.

[0157] Finally, based on the determined uncertainty score and diversity score of the image sample, the comprehensive score of the image sample can be determined. Specifically, refer to the following formula:

[0158] F(x) = U(X) + λD(X) (12)

[0159] In formula (12), F(x) is the comprehensive score of the image sample x, D(x) is the diversity score of the image sample x, λ is an adjustment parameter, and U(X) is the uncertainty score of the image sample X.

[0160] After that, from the unlabeled data Data(u), the top n samples with the comprehensive score are selected through the above image sample screening method, and the selected image samples are labeled and added to the labeled data Data(l). (Note: The labeled data here is the training data accumulated before.) The extended labeled data set is used to train the visual detection model, and based on this visual detection model, n samples are continuously selected from the remaining unlabeled data to iteratively train the visual detection model. Repeat this active learning process until the active learning termination condition is met. Among them, the active learning termination condition can be defined as the maximum labeled data volume being m. After k active learning iterations, the amount of screened and labeled data = k × n. Then the active learning termination condition is: when k × n ≥ m, stop the active learning iteration.

[0161] It can be understood that in the embodiments of the present application, data related to user information, etc. is involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards.

[0162] Next, continue to describe the exemplary structure of the software module implementation of the training device 555 of the visual detection model provided by the embodiments of the present application. In some embodiments, as Figure 2 shown, the software module stored in the training device 555 of the visual detection model in the memory 550 may include:

[0163] The detection module 5551 is used to perform class detection on a plurality of image samples collected by the floor sweeping robot through a visual detection model of the floor sweeping robot, and obtain the uncertainty scores of the image samples;

[0164] The determination module 5552 is used to extract features from each of the image samples through the visual detection model, obtain the image features of each of the image samples, and determine the diversity scores of the image samples based on the first similarity between the image features of the image samples;

[0165] The fusion module 5553 is used to fuse the uncertainty scores and diversity scores of the image samples to obtain the comprehensive scores of the image samples;

[0166] The training module 5554 is used to screen out the first image samples from the plurality of image samples based on the comprehensive scores, and train the visual detection model based on the first image samples.

[0167] In some embodiments, the above detection module 5551 is further used to perform class detection on each of the image samples collected by the floor sweeping robot, obtain a plurality of detection frames of the image sample, and one detection frame indicates the class to which a target object in the image sample belongs; determine the class uncertainty and confidence uncertainty of each detection frame; determine the uncertainty score of each detection frame based on the class uncertainty and confidence uncertainty of each detection frame; and average the uncertainty scores of each detection frame to obtain the uncertainty score of the image sample.

[0168] In some embodiments, the above detection module 5551 is further used to, for each of the detection frames, determine the total number of classes of the target objects included in the image sample based on the class to which the detection frame indicates the target object belongs, and the probability that the target object in the detection frame is predicted to be each of the classes; multiply the probability that the target object in the detection frame is predicted to be each of the classes by the logarithm of the corresponding probability to obtain the uncertainty that the target object in the detection frame is predicted to be each of the classes; and sum the uncertainties that the target object in the detection frame is predicted to be each of the classes to obtain the class uncertainty of the detection frame.

[0169] In some embodiments, the above detection module 5551 is further used to, for each of the detection frames, determine a first weight of the class uncertainty and a second weight of the confidence uncertainty; fuse the class uncertainty and the confidence uncertainty based on the first weight, the second weight, and the total number of classes to obtain the uncertainty score of the detection frame; wherein the uncertainty score is positively correlated with the class uncertainty and negatively correlated with the total number of classes and the confidence uncertainty.

[0170] In some embodiments, the above-mentioned determination module 5552 is further configured to, for each of the image samples, determine a first similarity between the image features of the image sample and the image features of other images; determine the difference between a preset value and each of the first similarities, and sum up each of the differences to obtain the diversity score of the image sample.

[0171] In some embodiments, the above-mentioned training module 5554 is further configured to sort the plurality of image samples in ascending order according to the level of the comprehensive score to obtain an image sample sequence; screen out a first number of image samples from the front end of the image sample sequence, and determine the screened first number of image samples as the first image samples.

[0172] In some embodiments, the above-mentioned training module 5554 is further configured to, based on the comprehensive score, screen out a second number of image samples from the plurality of image samples whose comprehensive score exceeds a score threshold, and determine the screened second number of image samples as the first image samples.

[0173] In some embodiments, when the number of the first image samples is multiple, the above-mentioned training module 5554 is further configured to determine a second similarity between each of the first image samples; screen out image sample pairs corresponding to the second similarity that exceeds a similarity threshold from the multiple first image samples; perform a deduplication process on the screened image sample pairs to obtain second image samples; and train the visual detection model based on the second image samples and third image samples in the first image samples other than the image sample pairs.

[0174] In some embodiments, the above-mentioned training module 5554 is further configured to, through a new visual detection model obtained by training the visual detection model based on the first image samples, perform class detection on fourth image samples in the plurality of image samples other than the first image samples to obtain uncertainty scores of each of the fourth image samples; extract image features of each of the fourth image samples through the new visual detection model, obtain the image features of each of the fourth image samples, and determine the diversity scores of each of the fourth image samples based on a third similarity between the image features of each of the fourth image samples; fuse the uncertainty scores and diversity scores of each of the fourth image samples to obtain the comprehensive scores of each of the fourth image samples; screen out fifth image samples from the fourth image samples based on the comprehensive scores, and perform iterative training on the new visual detection model based on the fifth image samples until the total number of image samples participating in the training reaches a quantity threshold and then stop the training.

[0175] An embodiment of the present application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer-executable instructions from the computer-readable storage medium, and the processor executes the computer-executable instructions, so that the electronic device executes the training method of the visual detection model described above in the embodiment of the present application.

[0176] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions or a computer program are stored. When the computer-executable instructions or the computer program are executed by a processor, it will cause the processor to execute the training method of the visual detection model provided by the embodiment of the present application. For example, as Figure 3A shown in the training method of the visual detection model.

[0177] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; it may also be various devices including one or any combination of the above memories.

[0178] In some embodiments, the computer-executable instructions may be in the form of a program, software, software module, script, or code, and may be written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including being deployed as an independent program or being deployed as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0179] As an example, the computer-executable instructions may or may not correspond to a file in the file system, and may be stored as a part of a file that stores other programs or data. For example, they may be stored in one or more scripts in a Hyper Text Markup Language (HTML) document, stored in a single file dedicated to the program being discussed, or stored in multiple cooperating files (for example, files that store one or more modules, subroutines, or code portions).

[0180] As an example, the computer-executable instructions may be deployed to execute on one electronic device, or on multiple electronic devices located at one location, or on multiple electronic devices distributed at multiple locations and interconnected by a communication network.

[0181] In summary, the following beneficial effects can be achieved through the embodiments of the present application:

[0182] Through the visual detection model of the floor cleaning robot, category detection is performed on multiple image samples collected by the floor cleaning robot to obtain the uncertainty scores of each image sample. By obtaining the uncertainty scores of the image samples through the visual detection model of the floor cleaning robot, a score for measuring the difficulty of performing category detection on the image samples can be obtained, thereby improving the accuracy of the comprehensive score obtained based on the uncertainty scores subsequently;

[0183] Through the visual detection model, feature extraction is performed on each image sample to obtain the image features of each image sample, and based on the first similarity between the image features of each image sample, the diversity scores of each image sample are determined. By obtaining the diversity scores of the image samples through the visual detection model of the floor cleaning robot, a score for measuring whether the image samples are duplicate samples can be obtained, thereby improving the accuracy of the comprehensive score obtained based on the diversity scores subsequently. Then, the uncertainty scores and diversity scores of each image sample are fused to obtain the comprehensive scores of each image sample. Based on the comprehensive scores, the first image samples are selected from multiple image samples, and the visual detection model is trained based on the first image samples. By using the comprehensive scores to represent the quality of the image samples, duplicate samples and high-quality first image samples that are easily recognizable can be eliminated, thereby improving the performance of the visual detection model trained based on the first image samples..

[0184] As described above, it is only the embodiment of the present application and is not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and scope of the present application are all included in the protection scope of the present application.

Claims

1. A method for training a visual detection model, characterized in that: The method comprises: Using a visual detection model of the sweeping robot, performing category detection on a plurality of image samples collected by the sweeping robot to obtain an uncertainty score for each of the image samples; Extracting features from each of the image samples using the visual detection model to obtain image features of each of the image samples, and determining a diversity score for each of the image samples based on a first similarity between the image features of each of the image samples; The uncertainty score and the diversity score of each of the image samples are integrated to obtain a comprehensive score of each of the image samples; A first image sample is screened out from the multiple image samples based on the comprehensive score, and the visual detection model is trained based on the first image sample.

2. The method according to claim 1, characterized in that The performing category detection on the multiple image samples collected by the sweeping robot to obtain an uncertainty score for each of the image samples includes: For each of the image samples collected by the sweeping robot, performing category detection on the image sample to obtain a plurality of detection frames of the image sample, wherein one detection frame indicates a category to which a target object in the image sample belongs; Determining the category uncertainty and confidence uncertainty of each of the detection boxes; Determining an uncertainty score for each of the detection boxes based on the category uncertainty and the confidence uncertainty of each of the detection boxes; The uncertainty scores of the detection frames are averaged to obtain the uncertainty score of the image sample.

3. The method according to claim 2, characterized in that The determining of the category uncertainty of each of the detection frames comprises: For each of the detection frames, based on the category to which the target object indicated by the detection frame belongs, determining the total number of categories of the target objects included in the image sample and the probability that the target object in the detection frame is predicted to belong to each of the categories; Multiplying the probability that the target object in the detection frame is predicted to be each of the categories by the corresponding logarithm of the probability to obtain the uncertainty that the target object in the detection frame is predicted to be each of the categories; The uncertainty of each category predicted for the target object in the detection frame is summed to obtain the category uncertainty of the detection frame.

4. The method according to claim 3, characterized in that The determining the uncertainty score of each detection frame based on the category uncertainty and the confidence uncertainty of each detection frame includes: For each of the detection frames, determining a first weight of each of the category uncertainties and a second weight of the confidence uncertainty; Based on the first weight, the second weight and the total number of categories, the category uncertainty and the confidence uncertainty are integrated to obtain an uncertainty score of the detection box; The uncertainty score is positively correlated with the category uncertainty, and negatively correlated with the total number of categories and the confidence uncertainty.

5. The method according to claim 1, characterized in that The determining, based on the first similarity between the image features of the image samples, a diversity score of each of the image samples comprises: For each of the image samples, determining a first similarity between an image feature of the image sample and image features of other images; The difference between the preset value and each of the first similarities is determined, and each of the differences is summed to obtain a diversity score of the image sample.

6. The method according to claim 1, characterized in that The step of selecting a first image sample from the plurality of image samples based on the comprehensive score comprises: Arranging the plurality of image samples in ascending order according to the comprehensive scores to obtain an image sample sequence; A first number of image samples are screened out from the front end of the image sample sequence, and the first number of image samples screened out are determined as first image samples.

7. The method according to claim 1, characterized in that The step of selecting a first image sample from the plurality of image samples based on the comprehensive score comprises: Based on the comprehensive score, a second number of image samples whose comprehensive scores exceed a score threshold are screened out from the multiple image samples, and the screened out second number of image samples are determined as first image samples.

8. The method according to claim 1, characterized in that The training of the visual detection model based on the first image sample comprises: In response to the number of the first image samples being plural, determining a second similarity between the first image samples; Filter out image sample pairs corresponding to the second similarity that exceed a similarity threshold from the plurality of first image samples; Performing deduplication processing on the screened image sample pairs to obtain second image samples; The visual detection model is trained based on the second image sample and a third image sample in the first image sample excluding the image sample pair.

9. The method according to claim 1, characterized in that: After training the visual detection model based on the first image sample, the method further includes: Using a new visual detection model obtained by training the visual detection model based on the first image sample, performing category detection on fourth image samples other than the first image sample among the multiple image samples to obtain an uncertainty score of each of the fourth image samples; Performing feature extraction on each of the fourth image samples by using the new visual detection model to obtain image features of each of the fourth image samples, and determining a diversity score of each of the fourth image samples based on a third similarity between the image features of each of the fourth image samples; fusing the uncertainty score and the diversity score of each of the fourth image samples to obtain a comprehensive score of each of the fourth image samples; A fifth image sample is selected from the fourth image sample based on the comprehensive score, and the new visual detection model is iteratively trained based on the fifth image sample until the total number of image samples involved in the training reaches a quantity threshold and the training is stopped.

10. A training device for a visual detection model, characterized in that: The device comprises: A detection module, used to perform category detection on a plurality of image samples collected by the sweeping robot through a visual detection model of the sweeping robot, and obtain an uncertainty score of each of the image samples; a determination module, configured to perform feature extraction on each of the image samples through the visual detection model to obtain image features of each of the image samples, and determine a diversity score of each of the image samples based on a first similarity between the image features of each of the image samples; A fusion module, used to fuse the uncertainty score and the diversity score of each image sample to obtain a comprehensive score of each image sample; A training module is used to select a first image sample from the multiple image samples based on the comprehensive score, and train the visual detection model based on the first image sample.

11. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, used to implement the training method of the visual detection model described in any one of claims 1 to 9 when executing the computer executable instructions or computer program stored in the memory.

12. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the training method of the visual detection model described in any one of claims 1 to 9 is implemented.

13. A computer program product comprising computer executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the training method of the visual detection model described in any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Sample screening method, computer equipment, storage medium and program product

    CN120411670A