Model training method and device, target detection method and device, equipment and storage medium
By performing domain transformation processing on the source domain image and using the object detection model to extract features, the problem of low detection accuracy in the target domain is solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411932355.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-24
- Publication Date
- 2025-05-06
AI Technical Summary
The traditional object detection model reduces the accuracy of object detection in the target domain, mainly due to the lack of data from the target domain during training.
By performing domain transformation processing on the source domain image, the target domain image is generated, and the image features of the source domain and the target domain are extracted using the target detection model to be trained, and the loss function is constructed to update the model parameters.
It improves the accuracy and generalization ability of target detection in the target domain, reduces the distribution differences between the source domain and the target domain, and enhances the robustness of the model.
Smart Images

Figure CN119942261A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to artificial intelligence technology, and in particular to a model training method, target detection method, device, equipment and storage medium. Background Art
[0002] The field of computer vision aims to enable computers to interpret and understand visual information in images and videos like humans. Object detection is one of the core tasks in the field of computer vision, which involves identifying and locating target objects in images or video frames and determining their boundaries. However, traditional object detection models are trained using source domain images that are easy to collect. Due to the lack of target domain data during the training process, the accuracy of object detection is reduced in the target domain space. Summary of the invention
[0003] The embodiments of the present application provide a model training method, a target detection method, an apparatus, a device and a storage medium, which can improve the accuracy of target detection in a target domain.
[0004] The technical solution of the embodiment of the present application is implemented as follows:
[0005] The present application provides a model training method, the method comprising:
[0006] Perform domain transformation on the source domain image to obtain the target domain image;
[0007] The following processing is performed by the target detection model to be trained: extracting a first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image;
[0008] The target detection model to be trained is used to perform the following processing: extracting a second image feature of the target domain image, and performing target prediction processing on the second image feature to obtain a second detection result of the target domain image;
[0009] Constructing a first loss based on the first image feature and the second image feature, and constructing a second loss based on the first detection result and the second detection result;
[0010] The parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain a trained target detection model.
[0011] In the above scheme, the reflectivity prediction of the first image feature to obtain the first reflectivity includes: performing a first convolution process on the first image feature to obtain a first convolution feature; performing a second convolution process on the first convolution feature to obtain a second convolution feature; and activating the second convolution feature to obtain the first reflectivity.
[0012] In the above scheme, the target prediction processing is performed on the first image feature to obtain the first detection result of the source domain image, including: mapping processing is performed on the first image feature to obtain mapping features; upsampling processing is performed on the mapping features to obtain upsampled features; and deconvolution processing is performed on the upsampled features to obtain the first detection result of the source domain image.
[0013] The present invention provides a method for detecting a target, which includes:
[0014] The following processing is performed through the target detection model: extracting the fourth image feature of the input image, and performing target prediction processing on the fourth image feature to obtain a detection result; wherein the target detection model is trained by the model training method provided in an embodiment of the present application.
[0015] The present application provides a model training device, including:
[0016] A domain transformation module is used to perform domain transformation processing on the source domain image to obtain a target domain image;
[0017] The target prediction module is used to perform the following processing through the target detection model to be trained: extracting the first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image; performing the following processing through the target detection model to be trained: extracting the second image feature of the target domain image, and performing target prediction processing on the second image feature to obtain a second detection result of the target domain image;
[0018] A parameter updating module is used to construct a first loss based on the first image feature and the second image feature, and to construct a second loss based on the first detection result and the second detection result; and to update the parameters of the target detection model to be trained based on the first loss and the second loss to obtain a trained target detection model.
[0019] The present application provides a target detection device, including:
[0020] The target detection module is used to perform the following processing through the target detection model: extract the fourth image feature of the input image, and perform target prediction processing on the fourth image feature to obtain a detection result; wherein the target detection model is trained by the model training method provided in the embodiment of the present application.
[0021] An embodiment of the present application provides an electronic device, including:
[0022] A memory for storing computer executable instructions;
[0023] The processor is used to implement the model training method or target detection method provided in the embodiment of the present application when executing the computer executable instructions stored in the memory.
[0024] An embodiment of the present application provides a computer-readable storage medium storing a computer program or computer-executable instructions for implementing the model training method or target detection method provided in the embodiment of the present application when executed by a processor.
[0025] An embodiment of the present application provides a computer program product, including a computer program or computer executable instructions. When the computer program or computer executable instructions are executed by a processor, the model training method or target detection method provided in the embodiment of the present application is implemented.
[0026] The embodiments of the present application have the following beneficial effects:
[0027] The source domain image is subjected to domain transformation processing to obtain the target domain image. Thus, the target domain image is obtained through domain transformation processing, and there is no need to manually collect images in the target domain space, thereby improving the efficiency and accuracy of image collection. A first loss is constructed based on the first image feature and the second image feature, and a second loss is constructed based on the first detection result and the second detection result. The parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain the trained target detection model. Thus, through the extracted first image feature of the source domain image and the second image feature of the target domain image, the target detection model to be trained can learn the source domain and the target domain from a feature perspective, thereby reducing the distribution difference between the source domain and the target domain. At the same time, by constructing the second loss according to the first detection result and the second detection result, the target detection model to be trained is more robust when facing different detection results, thereby reducing the deviation that may be caused by a single detection result. By fusing the first loss and the second loss, the generalization ability of the trained target detection model in the target domain and the accuracy of target detection in the target domain are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 It is a schematic diagram of the architecture of the model training system provided in the embodiment of the present application;
[0029] Figure 2A is a first structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0030] Figure 2B is a second structural schematic diagram of an electronic device provided in an embodiment of the present application;
[0031] Figure 3A It is a first flow chart of the model training method provided in the embodiment of the present application;
[0032] Figure 3B It is a second flow chart of the model training method provided in the embodiment of the present application;
[0033] Figure 3C It is a third flow chart of the model training method provided in the embodiment of the present application;
[0034] Figure 3D It is a fourth flow chart of the model training method provided in the embodiment of the present application;
[0035] Figure 4 This is a flowchart of obstacle detection for a sweeping machine adapted to a dark light environment with zero data collection cost provided by an embodiment of the present application;
[0036] Figure 5A It is a schematic diagram of the first structure of the cyclic generative adversarial network model provided in an embodiment of the present application;
[0037] Figure 5B It is a second structural diagram of the cyclic generative adversarial network model provided in an embodiment of the present application;
[0038] Figure 6 It is a schematic diagram of a decomposed network model provided in an embodiment of the present application.
[0039] It should be pointed out that the above-mentioned "first" and "second" are only used to distinguish different solutions, and do not represent the degree of superiority or inferiority of the solutions or the priority in the implementation process. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0041] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0042] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0043] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0044] The relevant data collection and processing in the embodiments of this application should be strictly in accordance with the requirements of relevant national laws and regulations when applied in examples, obtain the informed consent or separate consent of the personal information subject, and carry out subsequent data use and processing within the scope of authorization of laws and regulations and the personal information subject.
[0045] Unless otherwise defined, all technical and scientific terms used in the embodiments of the present application have the same meanings as those commonly understood by those skilled in the art. The terms used in the embodiments of the present application are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0046] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0047] 1) A source domain image is an image containing specific visual features. The embodiments of the present application do not limit the source domain image. The source domain image can be an image collected from a conventional lighting environment, and the conventional lighting environment can be natural light, artificial light source, etc.
[0048] 2) The target domain image is an image that applies the style of the source domain image. While maintaining the content of the source domain image, the style of the source domain image is changed to match the style of the target domain. The style can be a specific combination of visual elements such as color, light and shadow, and composition, which is used to represent a specific visual effect. The embodiments of the present application do not limit the target domain image. The target domain image can be an image collected from a dark light environment.
[0049] In the related art, traditional obstacle detection algorithms for sweeping robots are usually trained under good lighting conditions. However, in low-light environments (such as at night or in low-light indoors), the detection accuracy of the model decreases significantly. To address the above problems, the embodiments of the present application provide a model training method, device, electronic device, computer-readable storage medium, and computer program product to improve the accuracy of target detection in the target domain.
[0050] The model training method described in the embodiments of the present application can be applied to various fields, such as target detection and anomaly detection in dark environments. That is, the model training method in the embodiments of the present application is not limited to a certain field.
[0051] The following describes an exemplary application of the electronic device provided in the embodiment of the present application. The device provided in the embodiment of the present application can be implemented as a terminal or a server. The following describes an exemplary application when the device is implemented as a server.
[0052] See also Figure 1 , Figure 1 It is an architectural diagram of the model training system 100 provided in an embodiment of the present application. To support a model training application, a terminal (terminal 400 is shown as an example) is connected to a server 200 via a network 300. The network 300 may be a wide area network or a local area network, or a combination of the two.
[0053] The terminal 400 is used to send the source domain image and the target detection model to be trained to the server 200 through the network 300. The server 200 is used to first perform domain transformation processing on the source domain image to obtain the target domain image. Secondly, the target detection model to be trained is used to perform the following processing: extracting the first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image, extracting the second image feature of the target domain image, and performing target prediction processing on the second image feature to obtain a second detection result of the target domain image, then constructing a first loss based on the first image feature and the second image feature, and constructing a second loss based on the first detection result and the second detection result, updating the parameters of the target detection model to be trained based on the first loss and the second loss to obtain the trained target detection model, and performing the following processing through the trained target detection model: extracting the fourth image feature of the input image, and performing target prediction processing on the fourth image feature to obtain a detection result, and returning the detection result to the terminal 400, and the terminal 400 displays the detection result through the graphical interface 410.
[0054] An example of terminal 400 performing model training is described below.
[0055] In some embodiments, the terminal 400 can independently complete the model training task. For example, the terminal 400 is used to first perform domain transformation processing on the source domain image to obtain the target domain image. Secondly, the following processing is performed through the target detection model to be trained: extract the first image feature of the source domain image, and perform target prediction processing on the first image feature to obtain the first detection result of the source domain image, extract the second image feature of the target domain image, and perform target prediction processing on the second image feature to obtain the second detection result of the target domain image, then, construct a first loss based on the first image feature and the second image feature, and construct a second loss based on the first detection result and the second detection result, update the parameters of the target detection model to be trained based on the first loss and the second loss to obtain the trained target detection model, and perform the following processing through the trained target detection model: extract the fourth image feature of the input image, and perform target prediction processing on the fourth image feature to obtain the detection result, and display the detection result through the graphical interface 410.
[0056] In one implementation scenario, a server or a terminal may train a target detection model so that the trained target detection model performs target detection on an input dark-light image. First, domain transformation processing is performed on the conventional-light image to obtain a dark-light image. Secondly, the following processing is performed through the target detection model to be trained: a first image feature of the conventional-light image is extracted, and a target prediction processing is performed on the first image feature to obtain a first detection result of the conventional-light image. A second image feature of the dark-light image is extracted, and a target prediction processing is performed on the second image feature to obtain a second detection result of the dark-light image. Then, a first loss is constructed based on the first image feature and the second image feature, and a second loss is constructed based on the first detection result and the second detection result. The parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain a trained target detection model. The following processing is performed through the trained target detection model: a fourth image feature of the input dark-light image is extracted, and a target prediction processing is performed on the fourth image feature to obtain a detection result.
[0057] In one implementation scenario, a server or a terminal may train an anomaly detection model so that the trained anomaly detection model performs anomaly detection on an input abnormal image. First, an abnormal transformation process is performed on a normal image to obtain an abnormal image. Secondly, the following processes are performed through the anomaly detection model to be trained: a first image feature of the normal image is extracted, and an abnormality prediction process is performed on the first image feature to obtain a first detection result of the normal image. A second image feature of the abnormal image is extracted, and an abnormality prediction process is performed on the second image feature to obtain a second detection result of the abnormal image. Then, a first loss is constructed based on the first image feature and the second image feature, and a second loss is constructed based on the first detection result and the second detection result. The parameters of the anomaly detection model to be trained are updated based on the first loss and the second loss to obtain a trained anomaly detection model. The following processes are performed through the trained anomaly detection model: a fourth image feature of the input image is extracted, and an abnormality prediction process is performed on the fourth image feature to obtain a detection result.
[0058] In some embodiments, server 200 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), as well as big data and artificial intelligence platforms.
[0059] The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., but is not limited thereto. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application.
[0060] See also Figure 2A , Figure 2A is a first structural diagram of an electronic device provided in an embodiment of the present application, Figure 2A The electronic device 500 shown may be Figure 1 In the terminal 400 or server 200, the electronic device 500 includes: at least one processor 510, a memory 550, and at least one network interface 520. The various components in the server 200 are coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 2A Various buses are labeled as bus system 540 .
[0061] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0062] The user interface 530 includes one or more output devices 531 that enable presentation of media content, including one or more speakers and / or one or more visual display screens. The user interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls;
[0063] In some embodiments, when the terminal 400 independently completes the model training task, the server 200 provided in the embodiment of the present application does not include the user interface 530.
[0064] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, a hard disk drive, an optical disk drive, etc. The memory 550 may optionally include one or more storage devices that are physically located away from the processor 510.
[0065] The memory 550 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 550 described in the embodiments of the present application is intended to include any suitable type of memory.
[0066] In some embodiments, the memory 550 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0067] Operating system 551, including system programs for processing various basic system services and performing hardware-related tasks, such as framework layer, core library layer, driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0068] A network communication module 552, for reaching other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 include: Bluetooth, Wireless Compatibility Certification (WiFi), and Universal Serial Bus (USB);
[0069] a presentation module 553 for enabling presentation of information via one or more output devices 531 (e.g., display screen, speaker, etc.) associated with the user interface 530 (e.g., a user interface for operating peripherals and displaying content and information);
[0070] In some embodiments, when the model training task is completed independently by the terminal 400, the server 200 provided in the embodiment of the present application may not include the presentation module 553.
[0071] The input processing module 554 is used to detect one or more user inputs or interactions from one of the one or more input devices 532 and translate the detected inputs or interactions; in some embodiments, when the embodiment independently completes the model training task by the terminal 400, the server 200 provided in the embodiment of the present application may not include the presentation module 553.
[0072] In some embodiments, the model training device provided in the embodiments of the present application can be implemented in software. Figure 2A The model training device 555 stored in the memory 550 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: a domain transformation module 5551, a target prediction module 5552, and a parameter update module 5553. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be explained below.
[0073] In some embodiments, the target detection device provided in the embodiments of the present application can also be implemented in software, see Figure 2B , Figure 2B is a second structural diagram of an electronic device provided in an embodiment of the present application, Figure 2B Except for the target detection device 556 shown, the rest can be Figure 2A The target detection device 556 stored in the memory 550 may be software in the form of a program or a plug-in, including the following software modules: a target detection module 5561, which is logical and can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0074] It should be noted that in the model training example below, those skilled in the art can apply the target detection model trained by the model training method provided in the embodiment of the present application to target detection based on their understanding of the following text.
[0075] See also Figure 3A , Figure 3AThis is a first flow chart of the model training method provided in the embodiment of the present application, which will be combined with Figure 3A The steps shown are used to illustrate that the model training method provided in the embodiments of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following will illustrate the collaborative implementation by a server and a terminal as an example.
[0076] In step 101, domain transformation processing is performed on the source domain image to obtain a target domain image.
[0077] Here, domain transformation processing is a processing method for converting an image from a source domain space to a target domain space.
[0078] In some embodiments, see Figure 3B , Figure 3B is a second flow chart of the model training method provided in the embodiment of the present application, for Figure 3A Step 101 shown can be performed by Figure 3B Steps 1011 to 1012 are implemented as described in detail below.
[0079] In step 1011, the source domain image is encoded to obtain the encoding features of the source domain image.
[0080] Here, encoding is used to map the visualized data to a new feature space to obtain features for characterizing the potential structure of the source domain image. The embodiment of the present application does not limit the encoding method, and the encoding can be one-hot encoding, label encoding, etc.
[0081] In some embodiments, step 1011 can be implemented by: performing convolution processing on the source domain image to obtain convolution features, performing pooling processing on the convolution features to obtain pooling features, and performing mapping processing on the pooling features to obtain encoding features of the source domain image.
[0082] For example, for a source domain image (such as "[[(50, 60, 70), (120, 130, 140)], [(180, 190, 200), (220, 230, 240)]]", (50, 60, 70) is used to represent the pixel in the upper left corner of a red, green, and blue image containing 2*2 pixels, 50 is used to represent the intensity of the pixel in the red channel, 60 is used to represent the intensity of the pixel in the green channel, and 70 is used to represent the intensity of the pixel in the red channel. The convolution operation is performed on the pixel (to characterize the intensity of the pixel in the blue channel) to obtain the convolution feature (such as [0.14, 0.13, -0.06, 0.15, 0.03, 0.06]), the convolution feature is pooled to obtain the pooled feature (such as [0.14, 0.13, -0.06]), the pooled feature is mapped to obtain the encoding feature of the source domain image (such as [0.52, 0.13, -0.06, 0.15]).
[0083] In step 1012, domain features of the target domain sample are extracted, and based on the domain features, domain migration processing is performed on the encoding features of the source domain image to obtain the target domain image.
[0084] Here, the target domain sample is an image belonging to the target domain. The domain transfer process is used to apply the style of the target domain sample to the source domain image, thereby generating a target domain image belonging to the target domain.
[0085] In some embodiments, the above “extracting domain features of target domain samples” may be implemented by encoding the target domain samples to obtain domain features of the target domain samples. The above step of “encoding the target domain samples to obtain domain features of the target domain samples” is similar to the above step of “encoding the source domain images to obtain the encoded features of the source domain images”, and will not be repeated here.
[0086] In some embodiments, the step 1012 of "based on the domain features, performing domain migration processing on the coding features to obtain the target domain image" can be implemented in the following ways: extracting the mean and variance of the coding features, and extracting the mean and variance of the domain features; mapping the coding features based on the mean and variance of the coding features, and the mean and variance of the domain features to obtain migration features; decoding the migration features to obtain the target domain image.
[0087] It should be noted that the mapping process makes the source domain image have the style of the target domain sample by making the mean and variance of the coding feature consistent with the mean and variance of the domain feature, so that the generated target domain image and the target domain sample belong to the same domain. The mean of the coding feature is the average of multiple values of the coding feature. The variance of the coding feature is the average of the square of the difference between each value in the coding feature and the mean of the coding feature, which is used to characterize the discreteness of the data.
[0088] In some embodiments, the above-mentioned "extracting the mean and variance of the coding feature" can be achieved in the following manner: determining the sum of multiple eigenvalues in the coding feature as a first sum; determining the ratio of the first sum to the number of eigenvalues of the coding feature as the mean of the coding feature; determining the difference between each eigenvalue in the coding feature and the mean of the coding feature as a first difference; and determining the ratio of the sum of the squares of multiple first differences to the number of eigenvalues of the coding feature as the variance of the coding feature.
[0089] For example, given a coding feature of [2, 1, 3], the sum of multiple eigenvalues in the coding feature (such as 6) is determined as the first sum; the ratio (such as 2) of the first sum and the number of eigenvalues of the coding feature (such as 3) is determined as the mean of the coding feature; the difference between each eigenvalue in the coding feature and the mean of the coding feature is determined as the first difference (such as [0, -1, 1]); the ratio (such as 0.67) of the sum of the squares of multiple first differences (such as [0, 1, 1]) and the number of eigenvalues of the coding feature (such as 3) is determined as the variance of the coding feature.
[0090] Continuing from the above embodiment, the step of “extracting the mean and variance of the domain feature” is similar to the step of “extracting the mean and variance of the encoding feature”, which will not be repeated here.
[0091] Continuing from the above embodiment, the above “mapping the coding features based on the mean and variance of the coding features and the mean and variance of the domain features to obtain the migration features” can be implemented in the following way: performing the following processing for each eigenvalue of the coding feature: determining a first ratio of the difference between the eigenvalue and the mean of the coding feature to the variance of the coding feature, and determining the target eigenvalue as the sum of the product of the variance of the domain feature and the first ratio and the mean of the domain feature; combining multiple target eigenvalues to obtain the migration feature.
[0092] For example, given a coding feature of [2, 1, 3], take the last eigenvalue of the coding feature (such as 3) as an example to illustrate, determine the first ratio (such as 1.5) of the difference (such as 1) between the eigenvalue (such as 3) and the mean (such as 2) of the coding feature and the variance (such as 0.67) of the coding feature, and determine the product of the variance (such as 2) of the domain feature and the first ratio (such as 3) and the sum (such as 5) of the mean (such as 2) of the domain feature as the target eigenvalue; combine multiple target eigenvalues (such as [2, -1, 5]) to obtain the migration feature.
[0093] Continuing from the above embodiment, the above “decoding the migration features to obtain the target domain image” can be implemented in the following ways: mapping the migration features to obtain mapping features; upsampling the mapping features to obtain upsampled features; and deconvolving the upsampled features to obtain the target domain image.
[0094] For example, the migration features (such as [0.52, 0.13, -0.06, 0.15]) are mapped to obtain mapping features (such as [0.14, 0.13, -0.06]); the mapping features are upsampled to obtain upsampled features (such as [0.14, 0.13, -0.06, 0.15, 0.03, 0.06]); the upsampled features are deconvolved to obtain the target domain image (such as [[(50, 60, 70), (120, 130, 140)], [(180, 190, 200), (220, 230, 240)]]).
[0095] Through the embodiments of the present application, through the domain conversion processing method, the style of the source domain image can be converted into the style of the target domain, while the content of the source domain image remains unchanged, and the number of target domain images is increased, so that the target detection model to be trained can learn richer information from the target domain images.
[0096] Continue to see Figure 3A In step 102, the target detection model to be trained performs the following processing: extracting the first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image.
[0097] Here, the target detection model to be trained is a target detection model that has not been trained yet, and the target detection model to be trained is used to identify and locate one or more target objects in an image or video frame and the position of the target object. The first image feature is an attribute and information in the image that can be recognized and processed by an electronic device. The embodiment of the present application does not limit the first image feature. The first image feature can be a color feature, a texture feature, a shape feature, etc. of the image, wherein the color feature is used to describe the distribution and composition of the color in the image; the texture feature is used to describe the texture information of the image, such as roughness, smoothness, etc.; the shape feature is used to describe the shape of the object in the image, such as edges, contours, etc. The target prediction processing is used to map the features of the potential structure of the source domain image in the feature space into visual data. The first detection result is the result obtained by the target prediction processing. The embodiment of the present application does not limit the first detection result. The first detection result can be a detection frame of an object in the source domain image, a recognition result of the object in the detection frame, and the accuracy of the recognition result.
[0098] In some embodiments, the object detection model to be trained comprises N consecutive feature extraction layers, where N is a positive integer greater than 1, see Figure 3C , Figure 3C is a third flow chart of the model training method provided in the embodiment of the present application, for Figure 3A The first image feature of the source domain image extracted in step 102 can be obtained by Figure 3CSteps 1021 to 1023 are implemented as described in detail below.
[0099] In step 1021, feature extraction is performed on the source domain image through the first feature extraction layer to obtain image features output by the first feature extraction layer.
[0100] Here, the feature extraction layer is used to extract new features from the source domain image or feature to obtain deep features. The object detection model to be trained includes N consecutive feature extraction layers connected in series.
[0101] In some embodiments, step 1021 is similar to step 1011 and will not be described again herein.
[0102] In step 1022, the image features output by the n-1th feature extraction layer are extracted by the nth feature extraction layer to obtain the image features output by the nth feature extraction layer.
[0103] Wherein, n is a positive integer which increases successively, and 1<n≤N.
[0104] In some embodiments, step 1022 can be implemented in the following manner: performing convolution processing on the image features output by the n-1th feature extraction layer through the nth feature extraction layer to obtain the nth convolution feature, performing pooling processing on the nth convolution feature to obtain the nth pooling feature, and performing mapping processing on the nth pooling feature to obtain the image features output by the nth feature extraction layer.
[0105] Continuing from the above embodiment, the steps of "performing pooling processing on the nth convolution feature to obtain the nth pooling feature, performing mapping processing on the nth pooling feature, and obtaining the image feature output by the nth feature extraction layer" are similar to the steps of "performing pooling processing on the convolution feature to obtain the pooling feature, performing mapping processing on the pooling feature, and obtaining the encoding feature of the source domain image", which will not be repeated here.
[0106] Continuing from the above embodiment, the above “convolution processing of the image features output by the n-1th feature extraction layer” can be implemented in the following manner: determine the convolution kernel, perform sliding processing on the image features output by the n-1th feature extraction layer based on the convolution kernel, obtain multiple coverage areas, calculate the dot product of the convolution kernel and the image features output by the n-1th feature extraction layer in each coverage area, combine each dot product, and obtain the convolution feature, wherein the size of the coverage area is consistent with the size of the convolution kernel. It should be noted that the embodiment of the present application does not limit the convolution kernel, and the convolution kernel can be a Gaussian filter, a sharpening filter, etc. For example, a convolution kernel is determined (such as [1, 0, -1]), and based on the convolution kernel, the image features output by the n-1th feature extraction layer (such as [1, 2, 3, 4, 5, 4]) are slidingly processed to obtain multiple coverage areas (such as coverage area A is [1, 2, 3], coverage area B is [2, 3, 4], coverage area C is [3, 4, 5], and coverage area D is [4, 5, 4]). The dot product between the convolution kernel and the image features output by the n-1th feature extraction layer in each coverage area is calculated (such as the dot product of coverage area A is -2, the dot product of coverage area B is -2, the dot product of coverage area C is -2, and the dot product of coverage area D is 0), and each dot product is combined to obtain a convolution feature (such as [-2, -2, -2, 0]).
[0107] In step 1023, a first image feature is determined from the image features respectively output by the N feature extraction layers.
[0108] In some embodiments, step 1023 can determine the first image feature by one of the following methods: determining the image feature output by the target feature extraction layer as the first image feature, wherein the target feature extraction layer is a feature extraction layer that meets a preset serial number; when the dimension of any image feature is less than the first dimension threshold and the dimension of any image feature is greater than the second dimension threshold, determining any image feature as the first image feature, and the first dimension threshold is greater than the second dimension threshold; obtaining the execution resource amounts corresponding to N feature extraction layers respectively, and when the sum of the execution resource amounts of the current M feature extraction layers is greater than the available resources, determining the image feature output by the M-1th feature extraction layer as the first image feature, wherein M is a positive integer less than or equal to N.
[0109] Here, the preset sequence number is used to represent the order in which the feature extraction layer performs feature extraction.
[0110] It should be noted that the execution resource amounts corresponding to the N feature extraction layers are the amount of resources required by each feature extraction layer for feature extraction.
[0111] For example, the image feature output by the feature extraction layer that meets the preset sequence number (such as 3) is determined as the first image feature, that is, the image feature output by the third feature extraction layer is determined as the first image feature.
[0112] Continuing with the above example, when the dimension (such as 48*48) of any image feature (such as image feature A) is smaller than the first dimension threshold (such as 64*64), and the dimension of any image feature is larger than the second dimension threshold (such as 32*32), the any image feature is determined as the first image feature.
[0113] Continuing with the above example, given that the execution resource amount of the first feature extraction layer is 25, the execution resource amount of the second feature extraction layer is 20, and the execution resource amount of the third feature extraction layer is 15, when the sum of the execution resource amounts of the current three feature extraction layers (such as 60) is greater than the available resources (such as 55), the image feature output by the second feature extraction layer is determined as the first image feature. When the sum of the execution resource amounts of the first feature extraction layer (such as 25) is greater than the available resources (such as 20), other methods are used to determine the first image feature, where the other methods may be a method of determining the image feature output by the feature extraction layer that meets the preset sequence number as the first image feature, or a method of determining any image feature as the first image feature when the dimension of any image feature is less than the first dimension threshold and the dimension of any image feature is greater than the second dimension threshold, etc.
[0114] Through the embodiments of the present application, deep feature extraction is performed on the source domain image through continuous feature extraction layers to effectively capture key information in the image. The target detection model can more accurately identify and locate the target object in the image, thereby improving the accuracy and efficiency of target detection. By flexibly determining the first image feature, the performance and resource utilization of the target detection model can be optimized according to different application scenarios and resource constraints. The target detection model can optimize the feature extraction process according to the available resources and dimensionality thresholds, achieve optimal performance under limited computing resources, and improve the accuracy of target detection in the target domain. The number of continuous feature extraction layers N can be adjusted as needed to facilitate the expansion and optimization of the target detection model.
[0115] In some embodiments, the target prediction processing of the first image feature in step 102 to obtain the first detection result of the source domain image can be achieved by: mapping the first image feature to obtain the mapping feature; upsampling the mapping feature to obtain the upsampled feature; and deconvolution the upsampled feature to obtain the first detection result of the source domain image.
[0116] Here, the mapping process is used to map the fused features to a new feature space to obtain the mapping features of the potential structure used to characterize the spectral features. The embodiment of the present application does not limit the mapping process, and the mapping process can be linear mapping, nonlinear mapping, etc. The embodiment of the present application does not limit the upsampling process, and the upsampling process can be a nearest neighbor interpolation method, a bilinear interpolation method, etc. The upsampling process is used to increase the spatial dimension of the feature map. The nearest neighbor interpolation method is used as an example to illustrate that the nearest neighbor pixel points are selected to fill the new spatial position.
[0117] In some embodiments, taking the mapping process as a linear mapping process as an example, the above-mentioned "mapping the first image feature to obtain a mapping feature" can be achieved in the following way: determining a mapping parameter, and determining the product of the first image feature and the mapping parameter as the mapping feature.
[0118] For example, a mapping parameter (such as 2) is determined, and the product of the first image feature (such as [1, 0, -1]) and the mapping parameter (such as [2, 0, -2]) is determined as the mapping feature. The mapping feature (such as [2, 0, -2]) is upsampled to obtain an upsampled feature (such as [[1, 1, 2, 2], [1, 1, 2, 2], [3, 3, 4, 4], [3, 3, 4, 4]]).
[0119] Continuing from the above embodiment, the above-mentioned “performing deconvolution processing on the upsampled features to obtain the first detection result of the source domain image” can be implemented in the following way: determine the transposed convolution kernel, slide the transposed convolution kernel on the upsampled features by means of a sliding window, calculate the dot product of each element of the upsampled features and the transposed convolution kernel, and fuse each dot product to obtain the first detection result of the source domain image.
[0120] It should be noted that the deconvolution process is the reverse process of the convolution process, and the activation process introduces nonlinearity through the activation function to solve the nonlinear problem. The embodiment of the present application does not limit the activation process. The activation function can be a linear mapping unit (Rectified Linear Unit, ReLU), a leaky linear mapping unit (Leaky Rectified Linear Unit, Leaky-ReLU), etc., wherein the linear mapping unit is used to characterize that when the input is greater than 0, the input value is output, otherwise the output is 0. The linear mapping unit is simple to calculate and has a fast training speed.
[0121] For example, determine the transposed convolution kernel (such as [[1, 2], [3, 4]]), slide the transposed convolution kernel on the upsampled feature (such as [[1, 2], [3, 4]]) by means of a sliding window, calculate the dot product of each element of the upsampled feature and the transposed convolution kernel (such as the dot product corresponding to element 1 [[1, 2], [3, 4]], the dot product corresponding to element 2 [[2, 4], [6, 8]], the dot product corresponding to element 3 [[3, 6], [9, 12]], and the dot product corresponding to element 4 [[4, 8], [12, 16]]), and fuse each dot product to obtain the first detection result of the source domain image (such as [[1, 4, 4], [6, 20, 16], [9, 24, 16]]).
[0122] Continue to see Figure 3A In step 103, the target detection model to be trained performs the following processing: extracting the second image features of the target domain image, and performing target prediction processing on the second image features to obtain a second detection result of the target domain image.
[0123] In some embodiments, step 103 is similar to step 102 and will not be described in detail herein.
[0124] In step 104, a first loss is constructed based on the first image feature and the second image feature, and a second loss is constructed based on the first detection result and the second detection result.
[0125] Here, the second loss is an indicator for measuring the difference between the obtained detection result and the true label of the input image (for example, the difference between the first detection result and the target detection label of the source domain image, and the difference between the second detection result and the target detection label of the target domain image). The embodiment of the present application does not limit the second loss, and the second loss can be a mean square error loss, a cross entropy loss, etc. The first loss is used to allow the target detection model to be trained to learn an embedding space, so that similar samples are closer in this embedding space, and dissimilar samples are farther away in this embedding space, so that the intrinsic structure and regularity of the data can be better learned. The embodiment of the present application does not limit the construction method of the first loss, and the first loss can be the mean absolute error, mean square error, etc. of the first image feature and the second image feature.
[0126] In some embodiments, see Figure 3D , Figure 3D is a fourth flow chart of the model training method provided in the embodiment of the present application, for Figure 3A The first loss is constructed based on the first image feature and the second image feature in step 104, which can be obtained by Figure 3D Steps 1041 to 1044 are implemented as described in detail below.
[0127] In step 1041, the similarity between the first image feature and the second image feature is determined, and the difference between the preset similarity and the similarity is determined as the third loss.
[0128] For example, the similarity (e.g., 0.95) between a first image feature (e.g., [[(50, 60, 70), (120, 130, 140)], [(180, 190, 200), (220, 230, 240)]]) and a second image feature (e.g., [[(51, 61, 71), (120, 130, 140)], [(180, 190, 200), (220, 230, 240)]]) is determined, and the difference between a preset similarity (e.g., 1) and the similarity (e.g., 0.05) is determined as a third loss.
[0129] In step 1042, a reflectivity prediction is performed on the first image feature to obtain a first reflectivity, and a reflectivity prediction is performed on the second image feature to obtain a second reflectivity.
[0130] Here, reflectivity prediction is used to predict the reflectivity of an image (such as a source domain image and a target domain image). Reflectivity is used to characterize the degree to which an object in an image reflects light. For example, the reflectivity of a pure white pixel in an image is 1, which is used to characterize that the object corresponding to the pixel completely reflects the incident light; the reflectivity of a pure black pixel in an image is 0, which is used to characterize that the object corresponding to the pixel does not reflect the incident light at all.
[0131] In some embodiments, "predicting the reflectivity of the first image feature to obtain the first reflectivity" in step 1042 can be achieved by: performing a first convolution process on the first image feature to obtain a first convolution feature; performing a second convolution process on the first convolution feature to obtain a second convolution feature; and activating the second convolution feature to obtain the first reflectivity.
[0132] It should be noted that the activation processing introduces nonlinearity through the activation function to solve the nonlinear problem. The embodiment of the present application does not limit the activation processing. The activation function can be a linear mapping unit (Rectified Linear Unit, ReLU), a leaky linear mapping unit (Leaky Rectified Linear Unit, Leaky-ReLU), etc., wherein the linear mapping unit is used to characterize that when the input is greater than 0, the input value is output, otherwise the output is 0. The linear mapping unit is simple to calculate and has a fast training speed.
[0133] In some embodiments, the above-mentioned step of "performing a first convolution processing on the first image feature to obtain a first convolution feature" and the above-mentioned step of "performing a second convolution processing on the first convolution feature to obtain a second convolution feature" are similar to the above-mentioned step of "performing a convolution processing on the image feature output by the n-1th feature extraction layer", and will not be repeated here.
[0134] Continuing with the above embodiment, the first convolution processing can also be implemented through multiple dense convolution layers, and the input features of each layer are obtained by concatenating the output features of the previous layer and the input features of the previous layer, and the network structure of each dense convolution layer is consistent.
[0135] For example, taking three layers of dense convolutional layers as an example, the first image feature (such as [0.33, 0.24, 0, 0.25, 0.235]) is convolved through dense convolutional layer A to obtain a first dense convolutional feature (such as [0.34, 0.23, 0, 0.24, 0.22]), the first dense convolutional feature and the first image feature are concatenated to obtain a first concatenated feature (such as [0.33, 0.24, 0, 0.25, 0.235, 0.34, 0.23, 0, 0.24, 0.22]), and the first concatenated feature is concatenated through dense convolutional layer B. The second dense convolution feature is convolved with the first concatenated feature to obtain a second dense convolution feature such as [0.23, 0.3, 0, 0.2, 0.5]), the second dense convolution feature and the first concatenated feature are concatenated to obtain a second concatenated feature (such as [0.33, 0.24, 0, 0.25, 0.235, 0.34, 0.23, 0, 0.24, 0.22, 0.23, 0.3, 0, 0.2, 0.5]), and the second concatenated feature is convolved through a dense convolution layer C to obtain the sound feature of the first object (such as [0.32, 0.3, 0, 0.6, 0.1]).
[0136] In some embodiments, the above step of "predicting the reflectivity of the second image feature to obtain the second reflectivity" is similar to the above step of "predicting the reflectivity of the first image feature to obtain the first reflectivity", which will not be repeated here.
[0137] In step 1043, a fourth loss is constructed based on the first reflectivity and the second reflectivity.
[0138] In some embodiments, step 1043 can be implemented by: determining a first difference between a reflectivity label corresponding to a source domain image and a first reflectivity; determining a second difference between a reflectivity label corresponding to a target domain image and a second reflectivity; and fusing the first difference and the second difference to obtain a fourth loss.
[0139] In some embodiments, before the above “determining the first difference between the reflectivity label corresponding to the source domain image and the first reflectivity”, the following processing is performed: obtaining the reflectivity label corresponding to the source domain image. The above “obtaining the reflectivity label corresponding to the source domain image” can be implemented in the following manner: using a preset reflectivity model to predict the reflectivity of the source domain image, and obtaining the reflectivity label corresponding to the source domain image.
[0140] As an example, a preset reflectivity model including 5 feature extraction layers and 1 activation layer is used for explanation. The first feature extraction layer is used to extract features of the source domain image to obtain image features output by the first feature extraction layer. The nth feature extraction layer is used to extract features of the image features output by the n-1th feature extraction layer to obtain image features output by the nth feature extraction layer, wherein n is a successively increasing positive integer, 1<n≤5. The activation layer is used to activate the image features output by the 5th feature extraction layer to obtain a reflectivity label corresponding to the source domain image.
[0141] In some embodiments, the above-mentioned "fusing the first difference and the second difference to obtain the fourth loss" can be achieved by at least one of the following methods: performing a weighted summation of the absolute value of the first difference and the absolute value of the second difference to obtain the fourth loss; or performing a weighted summation of the square of the first difference and the square of the second difference to obtain the fourth loss.
[0142] For example, a first difference (such as 0.2) between a reflectivity label (such as 0.8) corresponding to a source domain image and a first reflectivity (such as 0.6) is determined; and a second difference (such as 0.5) between a reflectivity label (such as 0.8) corresponding to a target domain image and a second reflectivity (such as 0.3) is determined.
[0143] Continuing with the above example, the absolute value of the first difference (such as 0.2) and the absolute value of the second difference (such as 0.5) are weightedly summed (such as 0.35) to obtain the fourth loss; or, the square of the first difference (such as 0.04) and the square of the second difference (such as 0.25) are weightedly summed (such as 0.145) to obtain the fourth loss.
[0144] In step 1044, the third loss and the fourth loss are fused to obtain the first loss.
[0145] In some embodiments, step 1044 may be implemented in the following manner: performing a weighted summation on the third loss and the fourth loss to obtain the first loss.
[0146] For example, the third loss is 0.4 and the fourth loss is 0.8. When the third loss and the fourth loss have the same weight, that is, the weights of the third loss and the fourth loss are both 0.5, the first loss is 0.6. When the third loss and the fourth loss have different weights, that is, the weight of the third loss is 0.6 and the weight of the fourth loss is 0.4, the first loss is 0.56.
[0147] In some embodiments, "constructing a second loss based on the first detection result and the second detection result" in step 104 can be implemented by: determining a third difference between the detection result label of the source domain image and the first detection result; determining a fourth difference between the detection result label of the target domain image and the second detection result; and fusing the third difference and the fourth difference to obtain a second loss.
[0148] In some embodiments, the above-mentioned step of "determining the third difference between the detection result label of the source domain image and the first detection result; determining the fourth difference between the detection result label of the target domain image and the second detection result; fusing the third difference and the fourth difference to obtain the second loss" is similar to the above-mentioned step of "determining the first difference between the reflectivity label corresponding to the source domain image and the first reflectivity; determining the second difference between the reflectivity label corresponding to the target domain image and the second reflectivity; fusing the first difference and the second difference to obtain the fourth loss", which will not be repeated here.
[0149] Through the embodiments of the present application, by constructing the first loss, the embedding space learned by the target detection model can better distinguish between similar and dissimilar samples, and learn the illumination invariant features in the image, which helps to improve the generalization ability of the target detection model on the image features of different images. By constructing the second loss, the difference between the target detection result and the true label is measured, and the target object in the image can be more accurately identified and located. By predicting the reflectivity of the image and adapting to different lighting environments, the target detection model can learn the illumination invariant features in the image, thereby improving the interpretability and accuracy of the target detection model. By constructing multiple loss functions, the target detection model can generalize better during the training process, reducing the risk of overfitting.
[0150] Continue to see Figure 3A In step 105, the parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain the trained target detection model.
[0151] In some embodiments, step 105 can be implemented in the following manner: performing a weighted summation on the first loss and the second loss to obtain a target loss; and updating the parameters of the target detection model to be trained based on the target loss to obtain a trained target detection model.
[0152] Continuing from the above embodiment, the step of "taking a weighted sum of the first loss and the second loss to obtain the target loss" is similar to the step of "taking a weighted sum of the third loss and the fourth loss to obtain the first loss", which will not be repeated here.
[0153] Continuing from the above embodiment, the above-mentioned “updating the parameters of the target detection model to be trained based on the target loss to obtain the trained target detection model” can be implemented in the following way: updating the parameters of the target detection model to be trained based on the target loss until the target loss converges, and using the parameters of the target detection model to be trained when the target loss converges as the parameters of the trained target detection model.
[0154] Continuing from the above embodiment, the above “updating the parameters of the target detection model to be trained based on the target loss” can be implemented in the following way: performing backpropagation in the target detection model to be trained based on the target loss to obtain the gradient; and updating the parameters of the target detection model to be trained based on the gradient.
[0155] It should be noted that back propagation is implemented through the back propagation algorithm, and the gradient of the loss function to the parameters of the target detection model to be trained is calculated by the derivative chain rule.
[0156] In some embodiments, the target detection method provided in the embodiment of the present application will be described in conjunction with step 201. The target detection method provided in the embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following description will be made by taking the collaborative implementation of the server and the terminal as an example.
[0157] In step 201, the following processing is performed by the target detection model: extracting the fourth image feature of the input image, and performing target prediction processing on the fourth image feature to obtain a detection result.
[0158] Among them, the target detection model is trained by the model training method provided in the embodiment of the present application.
[0159] In some embodiments, step 201 is similar to step 102 and will not be described in detail herein.
[0160] Below, an exemplary application of the model training method provided in an embodiment of the present application in an actual application scenario will be described.
[0161] In the related technology, the traditional obstacle detection algorithm of sweepers is usually trained under good lighting conditions, but in low-light environments (such as at night or in low-light indoors), the detection accuracy of the model drops significantly. The difference in image quality caused by changes in lighting makes it difficult for the model to identify obstacles, and a large amount of low-light image data is required to train and verify the model. However, in low-light environments, data collection is not only difficult, but the labeling process is also complicated and time-consuming. The cost of data collection and labeling is extremely high, which is difficult to meet the needs of practical applications.
[0162] In order to solve the above problems, an embodiment of the present application proposes a model training method, which uses style transfer technology to generate dark light environment images, avoids the data collection process in the actual environment, reduces the data collection cost, and effectively learns the lighting-invariant features by combining feature alignment loss and estimated reflectivity, thereby significantly improving the detection accuracy of the model in low-light environments.
[0163] Take the low-light environment adaptive obstacle detection of a sweeper with zero data collection cost as an example, see Figure 4 , Figure 4 This is a flowchart of the obstacle detection of a sweeping machine adaptable to a dark-light environment with zero data collection cost provided by an embodiment of the present application. The process of detecting obstacles of a sweeping machine adaptable to a dark-light environment with zero data collection cost provided by an embodiment of the present application is explained below.
[0164] In step 401, a dark light image (ie, a target domain image) is generated.
[0165] Here, the image dataset collected by the sweeper under normal lighting conditions (i.e., normal lighting images) and the public low-light image dataset are divided into the source domain and the target domain. The Cycle-Consistent Adversarial Networks (CycleGAN) model is used to perform style transfer (i.e., domain transformation processing) to convert the normal lighting images in the source domain (i.e., source domain images) into low-light images similar to the target domain, thereby generating paired image data. Figure 5A , Figure 5A is a first structural diagram of a cyclic generative adversarial network model provided in an embodiment of the present application, such as Figure 5A As shown, the first generator is used to generate a target domain image A according to an input source domain image A, and the target domain image A and the target domain image B are input into the first discriminator for discrimination to obtain a first discrimination result, see Figure 5B , Figure 5B is a second structural diagram of the cyclic generative adversarial network model provided in the embodiment of the present application, such as Figure 5BAs shown, the second generator is used to generate a source domain image C according to the input target domain image C, and the source domain image C and the source domain image D are input into the second discriminator for discrimination to obtain a second discrimination result. After obtaining the first discrimination result and the first discrimination result, the parameters of the first generator, the second generator, the first discriminator and the second discriminator are updated according to the first discrimination result and the first discrimination result obtained by discrimination, so that the trained first generator generates the corresponding target domain image according to the input source domain image. The CycleGAN model can realize the style conversion of images under different lighting conditions without the need for paired training data. By training the generator and the discriminator, the style transfer from the conventional lighting image to the dark light image is completed. The image of the sweeper under conventional lighting conditions (i.e., the source domain image) is converted into a simulated dark light image through CycleGAN, which can effectively expand the training data at zero data collection cost and improve the detection ability of the model in a low-light environment.
[0166] In step 402, data is input into an object detection network.
[0167] Here, during the training process, the target detection network input is a set of image pairs, which are images collected in a normal lighting environment and corresponding dark light images generated based on the normal lighting images in the source domain. This set of images shares model parameters during the model training process. During the model testing process, only the image to be tested (i.e., the input image) needs to be input, where the image to be tested is an image without a label containing a real obstacle detection box.
[0168] In step 403, the object detection loss is calculated.
[0169] Here, during the model training process, the target detection prediction values of the two input images, that is, the obstacle detection boxes of the two input images (that is, the first detection result and the second detection result), are predicted based on the target detection network (for example, the yolov6 network model, that is, the target detection model to be trained), and the target detection loss L(Det) is calculated based on the obstacle detection boxes of the two input images as shown in Formula 1.
[0170]
[0171] Among them, x i is the input image, f(.) is the obstacle detection box output by the model, and y i It is the true value of the target detection corresponding to the input image, that is, the real detection box label corresponding to the input image.
[0172] In step 404, the feature alignment loss is calculated.
[0173] Here, during the model training process, since the two input images have different illumination but contain the same semantic information, by calculating the feature alignment loss L(Sim), the model can learn the illumination invariant features of the input image pair. The process of calculating the feature alignment loss is described in detail below.
[0174] First, the shallow features of the image pair (i.e., the first image features and the second image features) are extracted through the first few layers of the yolov6 model's neural network, and the deep features of the image pair are extracted through the yolov6 model's neural network. The deep features are mapped to obtain the target detection prediction values of the two input images, that is, the obstacle detection boxes of the two input images, where the shallow features include conventional lighting image features and dark light image features. The shallow features mainly include low-level lighting invariant features, such as edge and texture information. The shallow features remain consistent under conventional lighting and dark light environments. The yolov6 model includes multiple layers of neural networks.
[0175] Next, the cosine similarity is used to calculate the similarity between the regular light image features (i.e., the first image features) and the dark light image features (i.e., the second image features) to form the feature alignment loss L(Sim). By constructing the feature alignment loss, the model learns the low-level illumination invariant features. The calculation formula of the cosine similarity is shown in Formula 2.
[0176]
[0177] Among them, F n and F l They represent the features of the normal light image and the dark light image respectively, Sim(F n , F l ) is the similarity between the features of the regular light image and the dark light image, and the feature alignment loss is shown in Formula 3.
[0178] L(Sim)=1-Sim(F n , F l ) (3)
[0179] Among them, L(Sim) is the feature alignment loss.
[0180] In step 405, the reflectivity prediction loss is calculated.
[0181] Here, during the model training process, in order to further improve the model's ability to learn illumination-invariant features, a reflectivity decoding network is built to predict the reflectivity prediction value of the image based on shallow features, that is, to predict the reflectivity prediction value of the conventional illumination image. and the reflectance prediction value of the dark light image The reflectivity decoding network consists of two convolutional layers and one activation layer, which is responsible for decoding the input features (such as regular light image features and dark light image features) into reflectivity. The reflectivity decoding network is used to guide the generation of obstacle detection frames. Then, the decomposition network in the pre-trained image enhancement network (such as Retinext-Net) is used to estimate the reflectivity pseudo-true value of the two input images to obtain the reflectivity pseudo-true value R of the image pair. n and R l As a pseudo label, when dealing with the problem of obstacle detection in a dark light environment, in addition to considering the brightness change of the image, it is also necessary to pay attention to the reflectivity and illumination components in the image. Retinext-Net is an image decomposition network that can estimate the reflectivity and illumination information in the image, and generate pseudo labels of reflectivity for two images under different illumination conditions, thereby assisting the model to learn the illumination-invariant features, and using this as a supervisory signal to train the reflectivity decoding network, further enhancing the robustness of the obstacle detection model in a low-light environment. Figure 6 , Figure 6 is a schematic diagram of a decomposed network model provided in an embodiment of the present application, such as Figure 6 As shown, the decomposition network 602 is composed of multiple layers (e.g., 5 layers) of 3*3 convolutional layers with a step size of 1 and a padding of 1 and an activation layer, and is used to calculate the reflectivity pseudo-true value 603 (e.g., 0.8) of the input image according to the input image 601. Based on this, the reflectivity prediction loss L (ReF) is calculated as shown in Formula 4.
[0182]
[0183] in, is the predicted reflectance of the normal illumination image, is the predicted reflectance of the dark light image, R n is the reflectance pseudo-truth of the normal illumination image estimated using the pre-trained image enhancement network, R l is the reflectance true value of the dark light image estimated using the pre-trained image enhancement network, and L(ReF) is the reflectance prediction loss.
[0184] In step 406, the total loss is calculated.
[0185] Here, the total loss is used to train the model’s obstacle detection and illumination-invariant feature learning capabilities. The total loss function is shown in Formula 5.
[0186] L total =L(Det)+α*L(Sim)+β*L(ReF) (5)
[0187] Among them, α and β are the balancing factors of the multi-task loss function, which are used to adjust the weight of each task in the model parameter regression process.
[0188] In summary, the embodiment of the present application realizes the conversion from conventional lighting data to dark light environment images through the style transfer technology based on CycleGAN, avoiding the data collection process in the actual environment and reducing the data collection cost. By combining the reflectivity estimation and feature alignment mechanism of Retinext-Net, it is possible to effectively learn the features that are invariant to lighting, significantly improving the detection accuracy of the model in low-light environments.
[0189] The following is a description of an exemplary structure of a model training device 555 provided in an embodiment of the present application implemented as a software module. In some embodiments, Figure 2A As shown, the software modules stored in the model training device 555 of the memory 550 may include:
[0190] The domain transformation module 5551 is used to perform domain transformation processing on the source domain image to obtain a target domain image.
[0191] The target prediction module 5552 is used to perform the following processing through the target detection model to be trained: extract the first image feature of the source domain image, and perform target prediction processing on the first image feature to obtain a first detection result of the source domain image; perform the following processing through the target detection model to be trained: extract the second image feature of the target domain image, and perform target prediction processing on the second image feature to obtain a second detection result of the target domain image.
[0192] The parameter updating module 5553 is used to construct a first loss based on the first image feature and the second image feature, and to construct a second loss based on the first detection result and the second detection result; based on the first loss and the second loss, the parameters of the target detection model to be trained are updated to obtain the trained target detection model.
[0193] In some embodiments, the parameter update module 5553 is also used to determine the similarity between the first image feature and the second image feature, and determine the difference between the preset similarity and the similarity as the third loss; perform reflectivity prediction on the first image feature to obtain a first reflectivity, and perform reflectivity prediction on the second image feature to obtain a second reflectivity; construct a fourth loss based on the first reflectivity and the second reflectivity; and fuse the third loss and the fourth loss to obtain the first loss.
[0194] In some embodiments, the parameter updating module 5553 is also used to determine a first difference between a reflectivity label corresponding to a source domain image and a first reflectivity; determine a second difference between a reflectivity label corresponding to a target domain image and a second reflectivity; and fuse the first difference and the second difference to obtain a fourth loss.
[0195] In some embodiments, the target prediction module 5552 is also used to perform feature extraction on the source domain image through the first feature extraction layer to obtain image features output by the first feature extraction layer; perform feature extraction on the image features output by the n-1th feature extraction layer through the nth feature extraction layer to obtain image features output by the nth feature extraction layer, wherein n is a successively increasing positive integer, 1<n≤N; determine the first image feature from the image features output by the N feature extraction layers, and the target detection model to be trained includes N consecutive feature extraction layers, wherein N is a positive integer greater than 1.
[0196] In some embodiments, the target prediction module 5552 is also used to determine the first image feature in one of the following ways: determining the image feature output by the target feature extraction layer as the first image feature, wherein the target feature extraction layer is a feature extraction layer that meets a preset serial number; when the dimension of any image feature is less than the first dimension threshold and the dimension of any image feature is greater than the second dimension threshold, determining any image feature as the first image feature, and the first dimension threshold is greater than the second dimension threshold; obtaining the execution resource amounts corresponding to N feature extraction layers respectively, and when the sum of the execution resource amounts of the current M feature extraction layers is greater than the available resources, determining the image feature output by the M-1th feature extraction layer as the first image feature, wherein M is a positive integer less than or equal to N.
[0197] In some embodiments, the parameter updating module 5553 is further used to determine a third difference between the detection result label of the source domain image and the first detection result; determine a fourth difference between the detection result label of the target domain image and the second detection result; and fuse the third difference and the fourth difference to obtain a second loss.
[0198] In some embodiments, the domain transformation module 5551 is also used to encode the source domain image to obtain the encoding features of the source domain image; extract the domain features of the target domain sample, and based on the domain features, perform domain migration processing on the encoding features of the source domain image to obtain the target domain image.
[0199] In some embodiments, the domain transformation module 5551 is also used to extract the mean and variance of the coding features, and extract the mean and variance of the domain features; map the coding features based on the mean and variance of the coding features, and the mean and variance of the domain features to obtain migration features; decode the migration features to obtain the target domain image.
[0200] The following is a description of an exemplary structure of the target detection device 556 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2B As shown, the software modules stored in the target detection device 556 of the memory 550 may include:
[0201] The target detection module 5561 is used to perform the following processing through the target detection model: extract the fourth image feature of the input image, and perform target prediction processing on the fourth image feature to obtain a detection result; wherein the target detection model is trained by the model training method provided in the embodiment of the present application.
[0202] An embodiment of the present application provides a computer program product, which includes computer executable instructions. The computer executable instructions are stored in a computer-readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer-readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the above-mentioned model training method of the embodiment of the present application.
[0203] The present application embodiment provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions or computer programs are stored. When the computer-executable instructions or computer programs are executed by a processor, the processor will be caused to execute the model training method or target detection method provided by the present application embodiment, for example, FIG. 3A to FIG. 3D The model training method shown and Figure 4 The target detection method is shown.
[0204] In some embodiments, the computer-readable storage medium may be a memory such as RAM, ROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or may be various devices including one or any combination of the above memories.
[0205] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0206] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file that stores other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0207] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0208] In summary, the source domain image is subjected to domain transformation processing to obtain the target domain image. Thus, the target domain image is obtained through domain transformation processing, and there is no need to manually collect images in the target domain space, thereby improving the efficiency and accuracy of image collection. A first loss is constructed based on the first image feature and the second image feature, and a second loss is constructed based on the first detection result and the second detection result. The parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain the trained target detection model. Thus, through the extracted first image features of the source domain image and the second image features of the target domain image, the target detection model to be trained can learn the source domain and the target domain from a feature perspective, thereby reducing the distribution difference between the source domain and the target domain. At the same time, by constructing the second loss based on the first detection result and the second detection result, the target detection model to be trained is more robust in the face of different detection results, thereby reducing the deviation that may be caused by a single detection result. By fusing the first loss and the second loss, the generalization ability of the trained target detection model in the target domain and the accuracy of target detection in the target domain are improved.
[0209] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A model training method, characterized in that: The method comprises: Perform domain transformation on the source domain image to obtain the target domain image; The following processing is performed by the target detection model to be trained: extracting a first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image; The target detection model to be trained is used to perform the following processing: extracting a second image feature of the target domain image, and performing target prediction processing on the second image feature to obtain a second detection result of the target domain image; Constructing a first loss based on the first image feature and the second image feature, and constructing a second loss based on the first detection result and the second detection result; The parameters of the target detection model to be trained are updated based on the first loss and the second loss to obtain a trained target detection model.
2. The method according to claim 1, characterized in that The constructing a first loss based on the first image feature and the second image feature includes: Determining a similarity between the first image feature and the second image feature, and determining a difference between a preset similarity and the similarity as a third loss; Predicting the reflectivity of the first image feature to obtain a first reflectivity, and predicting the reflectivity of the second image feature to obtain a second reflectivity; constructing a fourth loss based on the first reflectivity and the second reflectivity; The first loss is obtained by fusing the third loss and the fourth loss.
3. The method according to claim 2, characterized in that The constructing a fourth loss based on the first reflectivity and the second reflectivity comprises: Determine a first difference between a reflectivity label corresponding to the source domain image and the first reflectivity; Determine a second difference between the reflectivity label corresponding to the target domain image and the second reflectivity; The first difference and the second difference are combined to obtain the fourth loss.
4. The method according to claim 1, characterized in that The target detection model to be trained comprises N consecutive feature extraction layers, where N is a positive integer greater than 1; The extracting a first image feature of the source domain image includes: Performing feature extraction on the source domain image through a first feature extraction layer to obtain image features output by the first feature extraction layer; The image features output by the n-1th feature extraction layer are extracted by the nth feature extraction layer to obtain the image features output by the nth feature extraction layer, wherein n is a positive integer that increases successively, and 1<n≤N; The first image feature is determined from the image features respectively output by the N feature extraction layers.
5. The method according to claim 4, characterized in that The determining the first image feature from the image features respectively output by the N feature extraction layers comprises: The first image feature is determined by one of the following methods: Determine the image feature output by the target feature extraction layer as the first image feature, wherein the target feature extraction layer is the feature extraction layer that meets the preset sequence number; When the dimension of any of the image features is smaller than a first dimension threshold and the dimension of any of the image features is larger than a second dimension threshold, any of the image features is determined as the first image feature, and the first dimension threshold is larger than the second dimension threshold; Obtain the execution resource amounts corresponding to the N feature extraction layers respectively. When the sum of the execution resource amounts of the current M feature extraction layers is greater than the available resources, determine the image feature output by the M-1th feature extraction layer as the first image feature, where M is a positive integer less than or equal to N.
6. The method according to any one of claims 1 to 5, characterized in that: The constructing a second loss based on the first detection result and the second detection result includes: Determine a third difference between the detection result label of the source domain image and the first detection result; Determine a fourth difference between the detection result label of the target domain image and the second detection result; The third difference and the fourth difference are combined to obtain the second loss.
7. The method according to claim 1, characterized in that The performing domain transformation processing on the source domain image to obtain the target domain image includes: Encoding the source domain image to obtain encoding features of the source domain image; Domain features of the target domain sample are extracted, and based on the domain features, domain migration processing is performed on the encoding features of the source domain image to obtain the target domain image.
8. The method according to claim 7, characterized in that The performing domain migration processing on the encoding feature based on the domain feature to obtain the target domain image includes: Extracting the mean and variance of the encoding feature, and extracting the mean and variance of the domain feature; Mapping the encoding feature based on the mean and variance of the encoding feature and the mean and variance of the domain feature to obtain a migration feature; The migration features are decoded to obtain the target domain image.
9. A target detection method, characterized in that: The method comprises: The target detection model is used to perform the following processing: extracting a fourth image feature of the input image, and performing target prediction processing on the fourth image feature to obtain a detection result; Wherein, the target detection model is trained by the model training method described in any one of claims 1 to 8.
10. A model training device, characterized in that: The device comprises: A domain transformation module is used to perform domain transformation processing on the source domain image to obtain a target domain image; The target prediction module is used to perform the following processing through the target detection model to be trained: extracting the first image feature of the source domain image, and performing target prediction processing on the first image feature to obtain a first detection result of the source domain image; performing the following processing through the target detection model to be trained: extracting the second image feature of the target domain image, and performing target prediction processing on the second image feature to obtain a second detection result of the target domain image; A parameter updating module is used to construct a first loss based on the first image feature and the second image feature, and to construct a second loss based on the first detection result and the second detection result; and to update the parameters of the target detection model to be trained based on the first loss and the second loss to obtain a trained target detection model.
11. A target detection device, characterized in that: The device comprises: A target detection module, used to perform the following processing through a target detection model: extracting a fourth image feature of an input image, and performing target prediction processing on the fourth image feature to obtain a detection result; wherein the target detection model is trained by the model training method described in any one of claims 1 to 8.
12. An electronic device, characterized in that: The electronic device comprises: A memory for storing computer executable instructions; The processor is used to implement the model training method described in any one of claims 1 to 8, or implement the target detection method described in claim 9 when executing the computer executable instructions or computer programs stored in the memory.
13. A computer-readable storage medium storing computer-executable instructions or a computer program, characterized in that: When the computer executable instructions or computer program are executed by a processor, the model training method described in any one of claims 1 to 8 is implemented, or the target detection method described in claim 9 is implemented.
Citation Information
Cited By
Target detection method and device, electronic equipment and computer program product
CN120912937A