Model training method and related device

By acquiring the sample set and using prior knowledge for inspection model training, the problems of inefficient and insufficient accuracy of inspection model update training are solved, and more efficient and accurate inspection model training is achieved.

WO2025130495A1PCT designated stage expired Publication Date: 2025-06-26HUAWEI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133604
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-11-21
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

The update training of the inspection model is inefficient and cannot improve the accuracy of the inspection model.

Method used

By obtaining the sample set, including multiple inspection-related pictures and their corresponding inspection item labels and results, model training is carried out based on this, and a priori knowledge assists model training is used to determine the training strategies for targeted inspection scenarios, and improve the efficiency and accuracy of model training.

Benefits of technology

It effectively improves the efficiency of model training and the accuracy of the inspection model, and can update and improve the inspection model more quickly, adapting to different inspection scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133604_26062025_PF_FP_ABST
    Figure CN2024133604_26062025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the field of artificial intelligence, and provide a model training method and a related device. The method comprises: first determining an inspection scenario on the basis of a plurality of first pictures in a sample set, and performing model training on the basis of the sample set and a first inspection model corresponding to the inspection scenario to obtain a second inspection model. The first inspection model is a preliminarily trained inspection model, and therefore, less training time is required for obtaining the second inspection model by training on the basis of the first inspection model, thereby improving the inspection model training efficiency; in addition, the inspection scenario is determined on the basis of the plurality of first pictures, and then the first inspection model is determined, such that training of inspection models of different inspection scenarios can be carried out in a targeted manner, thereby improving the precision of the inspection models.
Need to check novelty before this filing date? Find Prior Art

Description

A model training method and related equipment

[0001] This application claims priority to the Chinese patent application with application number 202311762679.3 filed with the State Intellectual Property Office of China on December 20, 2023, and priority to the Chinese patent application with the invention name “A model training method and related equipment”, all contents of which are incorporated by reference into this application. Technical Field

[0002] The embodiments of the present application relate to the field of artificial intelligence, and in particular to a model training method and related equipment. Background Art

[0003] Artificial Intelligence (AI) is the theory, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a branch of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. Research in the field of AI includes robotics, natural language processing, computer vision, decision-making and reasoning, human-computer interaction, recommendation and search, and basic AI theory.

[0004] Computer vision technology can be applied to patrol inspections. With the advancement of digital and intelligent patrol inspections, intelligent patrol inspection models are being widely used. Using patrol inspection models to test and review patrol items reduces the experience requirements for patrol inspection personnel and improves the efficiency of patrol inspections and subsequent manual review.

[0005] However, updating the inspection model not only has low training efficiency but also fails to improve the accuracy of the inspection model.

[0006] Therefore, how to solve the above problems has become a hot topic being studied by those skilled in the art. Summary of the Invention

[0007] This application provides a model training method and related equipment, which can effectively improve the efficiency of model training and the accuracy of inspection models.

[0008] In a first aspect, a model training method is provided, which can be executed by a model training device or a chip in the model training device.

[0009] The model training method includes the following steps: first, obtaining a sample set. The sample set includes multiple first images, as well as inspection item labels and inspection results for each of the multiple first images. Model training is performed based on the sample set and the first inspection model to obtain a second inspection model; the first inspection model is used to inspect at least one first inspection item in an inspection scenario. The inspection scenario is determined based on the multiple first images. The at least one first inspection item intersects with the inspection item corresponding to the sample set.

[0010] The first image is an image related to the inspection site, and the first inspection model and the second inspection model are used for inspection processing of the inspection site. The inspection item label is used to indicate the category of the inspection item in the first image, and the inspection result is used to indicate the inspection result corresponding to the inspection item in the first image.

[0011] As can be seen, in this solution, after first determining the inspection scene based on multiple first images, model training is performed based on the sample set and the first inspection model corresponding to the above inspection scene to obtain the second inspection model. Since the first inspection model is a preliminarily trained inspection model, the training time required to obtain the second inspection model based on the first inspection model is relatively short, which can improve the efficiency of inspection model training. In addition, determining the inspection scene based on multiple first images and then determining the first inspection model can enable targeted training of inspection models for different inspection scenes, which helps to improve the accuracy of the inspection model.

[0012] In one possible implementation of the first aspect, the model training based on the sample set and the first inspection model to obtain the second inspection model specifically includes the following steps: determining a correspondence between an inspection item label and the at least one first inspection item based on the sample set and a baseline inspection image for each of the at least one first inspection item; determining a training strategy for the first inspection model based on the correspondence; and training the first inspection model based on the sample set and the training strategy to obtain the second inspection model.

[0013] In this solution, since the inspection item label may be different from the name of the first inspection item, the baseline inspection picture of the first inspection item and the first picture in the sample set can be used to determine whether there is a corresponding relationship between the inspection item label and the first inspection item. Determining whether the inspection item label and the first inspection item are similar (i.e., corresponding) based on the picture makes the judgment result more accurate and reliable. The embodiment of the present application then determines the training strategy of the first inspection model based on the above-mentioned corresponding relationship, so as to perform model training on the first inspection model according to the above-mentioned training strategy and sample set to obtain the second inspection model.

[0014] In a possible implementation of the first aspect, the above-mentioned determination of the correspondence between the inspection item label and at least one first inspection item based on the sample set and the benchmark image of at least one first inspection item specifically includes the following steps: using a feature extraction network to extract features from the benchmark inspection image of the second inspection item to obtain a first benchmark feature. At least one first inspection item includes a second inspection item. Based on the first benchmark feature and the first feature of each inspection item label in the sample set, determine the first similarity between each inspection item label and the second inspection item. The first feature of the inspection item label is determined based on the image area where the inspection item is located in the first image. Determining the above-mentioned correspondence includes that the inspection item label with the highest first similarity among the inspection item labels corresponds to the second inspection item. Wherein, the benchmark inspection image is a picture including the above-mentioned second inspection item.

[0015] In this solution, a first similarity is determined between the first baseline feature and the first feature of the inspection item label, and the inspection item label with the highest first similarity is assigned to the second inspection item. This determines the correspondence between the inspection item label and the first inspection item, facilitating the determination of the training strategy for the first inspection model. This embodiment of the present application performs feature extraction on an image to obtain representative features, allowing for more accurate acquisition of the first similarity. Furthermore, based on the first similarity, it is possible to accurately determine whether the inspection item label and the second inspection item correspond.

[0016] In a possible implementation of the first aspect, the model training method further includes the following steps: determining k first images corresponding to each inspection item label in the sample set, where k is greater than or equal to 1. Cropping k second images of the area where the inspection items are located from the k first images of each inspection item label. Using a feature extraction network, performing feature extraction on the k second images of each inspection item label to obtain k second features of each inspection item label. Fusion of the k second features of each inspection item label to obtain the first feature of each inspection item label.

[0017] In this solution, the second image is an image that includes the inspection item. The second image can be the first image itself or a localized image of the first image. After feature extraction and feature fusion of the k second images for each inspection item label, the first feature of each inspection item label can be obtained. The first feature corresponding to the inspection item label is obtained based on the k first images of the inspection item label, which has a higher accuracy.

[0018] In a possible implementation of the first aspect, the training strategy for determining the first inspection model based on the correspondence specifically includes the following steps: determining at least one third inspection item of the first inspection model based on the correspondence; the at least one third inspection item is an inspection item that requires training of the first inspection model. When at least one third inspection item corresponds one-to-one with at least one first inspection item, the training strategy is to load the model parameters of the first inspection model and then perform model training. When at least one third inspection item corresponds one-to-one with part of the inspection items of at least one first inspection item, the training strategy is to load the model parameters of the first inspection model, keep the parameters of the backbone network in the first inspection model unchanged, and perform model training on the detection head in the first inspection model. When part of the inspection items of at least one third inspection item corresponds one-to-one with part of the inspection items of at least one first inspection item, the training strategy is to load the parameters of the backbone network in the first inspection model, and randomly initialize the parameters of the detection head in the first inspection model and then perform model training.

[0019] In this solution, after determining the correspondence between the first inspection item and the inspection item label, at least one third inspection item of the first inspection model can be determined based on the correspondence. The third inspection item can be obtained from the inspection items corresponding to the first inspection item and the inspection item label. The training strategy for the first inspection model is then determined based on the correspondence between the third inspection item and the first inspection item. Different model training strategies can be determined to suit the specific circumstances of the third inspection item, thereby improving the speed of model training and the accuracy of the second inspection model obtained after training.

[0020] In a possible implementation of the first aspect, when the number of first images corresponding to each inspection item in some or all of the at least one third inspection item is less than or equal to a quantity threshold, the first inspection model is a small-sample inspection model. A small-sample inspection model refers to an inspection model that can be trained using a relatively small number of samples.

[0021] Exemplarily, the small sample inspection model includes a prototype network module.

[0022] In this solution, when the number of sample images of all or part of the third inspection items is small, the first inspection model can adopt a small sample inspection model to successfully complete the model training to obtain the second inspection model.

[0023] In a possible implementation of the first aspect, the model training method further includes the following steps: determining the inspection scenes corresponding to some or all of the first pictures in the plurality of first pictures, and taking the inspection scene with the largest number of pictures as the inspection scene corresponding to the sample set.

[0024] In this solution, a voting mechanism is used to determine the inspection scenario corresponding to the sample set to ensure the accuracy of the inspection scenario.

[0025] In a second aspect, the present application also provides a model training device, which includes a unit or module for executing the model training method of the first aspect.

[0026] In a third aspect, the present application also provides a model training device, which includes a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program code, and the processor is used to call the program code to execute the model training method described in the first aspect.

[0027] In a fourth aspect, the present application also provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the model training method as described in the first aspect.

[0028] In a fifth aspect, the present application also provides a computer program product comprising instructions, which, when run on a computer, enables the computer to execute the model training method described in the first aspect.

[0029] In the sixth aspect, the present application also provides a chip, which includes a processor and a data interface. The processor reads instructions stored in the memory through the data interface to execute the model training method described in the first aspect.

[0030] Optionally, as an implementation method, the chip may further include a memory, in which instructions are stored, and the processor is used to execute the instructions stored on the memory. When the instructions are executed, the processor is used to execute the model training method described in the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The following is an introduction to the drawings used in the embodiments of this application.

[0032] FIG1A is a schematic diagram of a system framework of a model training method provided in an embodiment of the present application;

[0033] FIG1B is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0034] FIG2 is a flow chart of a model training method provided in an embodiment of the present application;

[0035] FIG3A is a schematic diagram of a specific flow chart of step 202 provided in an embodiment of the present application;

[0036] FIG3B is a schematic diagram of determining a third detection item provided by an embodiment of the present application;

[0037] FIG4 is a schematic structural diagram of a model training device provided in an embodiment of the present application;

[0038] FIG5 is a schematic structural diagram of another model training device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0039] The technical solution in this application will be described below with reference to the accompanying drawings.

[0040] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described in this application as "exemplary" or "for example" should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0041] The "at least one" mentioned in the embodiments of this application refers to one or more, and "plurality" refers to two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can be represented by: a, b, c, (a and b), (a and c), (b and c), or (a and b and c), where a, b, c can be single or multiple. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can be represented by: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship. The serial numbers of the steps in the embodiments of this application (such as step S1, step S21, etc.) are only for distinguishing different steps and do not limit the order of execution between the steps.

[0042] Furthermore, unless otherwise specified, ordinal numbers such as "first" and "second" in the embodiments of this application are used to distinguish multiple objects and are not used to limit the order, timing, priority, or importance of multiple objects. For example, the first device and the second device are only for ease of description and do not indicate differences in structure, importance, etc. between the first and second devices. In some embodiments, the first device and the second device can also be the same device.

[0043] In the above embodiments, the term "when" can be interpreted to mean "if...", "after...", "in response to determining...", or "in response to detecting...", depending on the context. The above are merely optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the concepts and principles of the present application shall be included in the scope of protection of the present application.

[0044] To facilitate understanding, the following first introduces relevant terms and other related concepts involved in the embodiments of this application.

[0045] (1) Inspection model

[0046] The inspection model can be used to perform inspections on a target inspection site. For example, it can detect the presence of inspection items within the target inspection site. Optionally, the operating status of the inspection item can also be detected, indicating whether the operating status is normal or abnormal. For example, the inspection model can be understood as a target detection model.

[0047] In the embodiments of this application, the inspection model is applicable to various scenarios, including but not limited to the following areas:

[0048] Power industry: Inspection models are used to monitor and detect power generation equipment, transmission equipment, and distribution equipment in real time, improving the safety and stability of the power system.

[0049] Manufacturing: Inspection models are used to monitor and identify equipment failures and anomalies on production lines, enabling proactive repairs and maintenance to ensure production continuity and stable quality.

[0050] Construction industry: Inspection models are used for building safety monitoring and equipment management, real-time monitoring of safety hazards such as fires and gas leaks, and providing timely alarms and treatment suggestions.

[0051] Transportation industry: Inspection models are used to monitor and manage traffic lights, tunnel equipment, traffic monitoring systems, etc., providing real-time traffic conditions and fault information, helping traffic management departments take timely measures to improve the efficiency and safety of transportation.

[0052] Fixed line inspection: such as power inspection, oil well inspection, bridge inspection, etc.

[0053] Patrols in fixed areas: such as public security patrols, park inspections, river inspections, reservoir patrols, etc.

[0054] General surveying and mapping scenarios: such as natural resource surveys, smart construction sites, and construction supervision.

[0055] (2) Inspection site

[0056] The inspection site is the target location for inspection and can be any location requiring inspection, such as a computer room, base station, or construction site. When the inspection site is a computer room, inspection items may include air conditioners and AC distribution boxes. When the inspection site is a base station, inspection items may include base station main equipment, base station AC / DC power distribution equipment, base station batteries, and base station antenna feeder systems. When the inspection site is a construction site, inspection items may include safety wearables such as hard hats and safety warning clothing.

[0057] Exemplarily, the inspection site may also be an institution or a park, etc. Taking an institution as an example, the institution may include multiple inspection scenes such as a computer room, an outdoor antenna, etc.

[0058] (3) Small sample learning model

[0059] A small sample learning model is a machine learning method that aims to build an accurate machine learning model using less training data. The specific number of samples in a small sample learning model can be determined based on the actual situation.

[0060] (4) Prototypical Networks

[0061] A prototypical network is a small-sample learning model that projects samples onto a metric space, placing similar samples closer together and heterogeneous samples farther apart. The prototypical network exploits this property to learn feature representations of samples and uses these representations for classification and prediction.

[0062] With the development of digital and intelligent inspection operations, intelligent inspection models have been widely used in inspection operations. By using inspection models to detect and review inspection items, the experience requirements for inspection personnel can be reduced, while also improving the efficiency of inspection implementation and subsequent manual review. However, updating the inspection model not only has low training efficiency, but also fails to improve the accuracy of the inspection model. Therefore, the embodiments of the present application provide a model training method that can effectively improve the efficiency of model training and the accuracy of the inspection model.

[0063] The model training method in the embodiment of the present application can be executed by a model training device, or by a chip in the model training device.

[0064] The following introduces the system framework of the model training method of the embodiment of the present application.

[0065] Refer to Figure 1A, which is a system framework diagram of a model training method provided in an embodiment of the present application; the application system framework of the model training method in an embodiment of the present application includes a work order configuration system, inspection equipment and model training equipment.

[0066] The work order configuration system is responsible for issuing inspection work orders to inspection equipment and receiving inspection results from the inspection equipment. The work order configuration system can issue inspection work orders in response to triggers from staff or third-party monitoring systems, without any specific restrictions.

[0067] Inspection devices receive inspection work orders from the work order configuration system, trigger inspection operations, and obtain inspection results. Inspection devices use the inspection model issued by the model training device to perform inspection operations.

[0068] Exemplarily, the inspection equipment scans the inspection site according to the work order, that is, records the video in real time, and obtains the picture to be processed by extracting single-frame images from the video at regular intervals. The inspection equipment loads the inspection model to process the picture to be processed to detect whether there is an inspection item. When there is an inspection item indicated by the work order in the picture to be processed, the inspection model generates an inspection result picture corresponding to the picture to be processed (for example, the inspection result picture includes the inspection item label and the inspection result), and transmits the inspection result picture to the model training device through the data recovery link. In addition, the inspection equipment generates an inspection result based on the inspection result picture, and feeds the inspection result back to the work order configuration system. Exemplarily, the inspection result can be marked on the inspection work order and then the marked inspection work order can be fed back to the work order configuration system. The work order configuration system can perform secondary applications such as statistics and analysis on the inspection results.

[0069] The model training device is used to execute the model training method of the embodiment of the present application and send the trained inspection model to the inspection device. The model training device can be a platform deployed locally or in the cloud for AI services such as image storage, annotation, model training, and model publishing.

[0070] Exemplarily, the model training device stores the received inspection result pictures, builds a sample set based on the inspection result pictures, and then determines the scene template required for training based on the sample set, and loads the corresponding preset model of the scene (for example, the first inspection model in the embodiment of the present application) and training configuration parameters according to the scene template. Then the model training device determines the training strategy of the preset model, and finally trains the preset model according to the sample set and the training strategy. After the training, the inspection model (for example, the second inspection model in the embodiment of the present application) is obtained. The inspection model is sent to the inspection device for model updating. Among them, the above-mentioned preset model includes at least one first inspection item in the scene (which can be understood as the default inspection item of the scene). In addition, when the number of sample pictures of the inspection items that need to be trained is small, the preset model can adopt a small sample inspection model to successfully complete the model training. The small sample inspection model is a small sample learning model.

[0071] The following describes a system architecture provided by an embodiment of the present application.

[0072] 1B , an embodiment of the present application provides a system architecture 100 . As shown in the system architecture 100 , a data acquisition device 160 is used to collect training data for the inspection model. In this embodiment, the training data is a plurality of first images. The data acquisition device 160 stores the training data in a database 130 .

[0073] The training device 120 can perform model training based on the training data maintained in the database 130 to obtain the inspection model 101 (such as the second inspection model in the embodiment of the present application). The training device 120 is the model training device of the embodiment of the present application. The specific model training process can refer to the specific description of the model training method (such as Figure 2) below, which will not be repeated here. For example, in the embodiment of the present application, after completing the model training, the inspection model 101 can obtain the ability to detect inspection items based on pictures. The training device 120 can be a server or a cloud service device, etc., and can also be a mobile phone terminal, a tablet computer, a laptop computer, AR / VR, a vehicle-mounted terminal, a monitoring device, a vehicle-mounted automatic driving system, a drone and other equipment.

[0074] It should be noted that, in actual applications, the training data maintained in the database 130 may not all be collected by the data acquisition device 160, but may also be received from other devices. It should also be noted that the training device 120 may not perform model training entirely based on the training data maintained by the database 130, but may also obtain training data from the cloud or other places for model training. The above description should not be used as a limitation on the embodiments of the present application.

[0075] The inspection model 101 obtained after processing by the training device 120 can be applied to the system or device deployed in the terminal, such as the terminal device 110 shown in Figure 1B. The terminal device 110 can process the input image based on the inspection model 101 to obtain the corresponding model processing results (including inspection item labels and inspection results of inspection items). The terminal device 110 is the inspection device in the embodiment of the present application. The terminal device 110 can be a terminal, such as a mobile phone terminal, a tablet computer, a laptop computer, AR / VR, a vehicle-mounted terminal, a vehicle-mounted computing platform, a vehicle-mounted domain controller, a monitoring device, a vehicle-mounted automatic driving system, a drone, etc. It can also be a server or a cloud device, etc. The training device 120 and the terminal device 110 can be the same device.

[0076] In FIG. 1B , the terminal device 110 is configured with an I / O interface 112 for interacting with external devices. The user can input data into the I / O interface 112 through the client device 140. In this embodiment, the input data is a picture of the inspection site, which can be input by the user or obtained from the database 130. The client device 140 can be a picture acquisition device, such as a mobile phone.

[0077] Optionally, the preprocessing module 113 is configured to preprocess the input data received by the I / O interface 112. In an embodiment of the present application, the preprocessing module 113 is configured to preprocess the input data received by the I / O interface 112, and the preprocessed data enters the computing module 111. In an embodiment of the present application, the input data may be an image, and the preprocessing module 113 may be configured to perform at least one preprocessing operation on the image, including scaling, normalization, flipping, rotation, cropping, expansion, and grayscale processing.

[0078] Scaling adjusts the image size to the network input layer size. Normalization normalizes the image pixel values ​​to a certain range, usually between 0 and 1. Flipping flips the image left-right or up-down. Rotating rotates the image by a certain angle. Cropping crops the image to the network input layer size. Dilation pads the image with additional pixels. Grayscaling converts a color image to grayscale.

[0079] When the terminal device 110 preprocesses the input data, or when the computing module 111 of the terminal device 110 performs calculations and other related processing, the terminal device 110 can call the data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.

[0080] Finally, the I / O interface 112 returns the model processing result of the input data to the client device 140 , thereby providing it to the user. In this case, the client device 140 may be a display.

[0081] Optionally, the model processing result can also be used as input to the calculation module, and the calculation module performs other processing operations based on the model processing result, such as obtaining the final inspection result based on the model processing result, and returning the inspection result to the client device 140 through the I / O interface 112.

[0082] In the scenario shown in FIG. 1B , the user can manually input data, which can be performed through the interface provided by I / O interface 112. Alternatively, client device 140 can automatically send input data to I / O interface 112. If user authorization is required for client device 140 to automatically send input data, the user can set the corresponding permissions in client device 140. The user can view the output of terminal device 110 on client device 140, which can be presented in a display, sound, action, or other specific form. Client device 140 can also serve as a data acquisition terminal, collecting input data and output results from I / O interface 112 as shown in FIG. 1B as new sample data and storing them in database 130. Of course, collection can also be performed without client device 140, with I / O interface 112 directly storing the input data and output results from I / O interface 112 as shown in FIG. 1B as new sample data in database 130.

[0083] It is worth noting that FIG1B is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationship between the devices, components, modules, etc. shown in FIG1B does not constitute any limitation. For example, in FIG1B, the data storage system 150 is an external memory relative to the terminal device 110. In other cases, the data storage system 150 can also be placed in the terminal device 110.

[0084] The model training method of the embodiment of the present application is described in detail below.

[0085] In the embodiment of the present application, the executor of the model training method takes the model training device as an example.

[0086] Referring to FIG. 2 , FIG. 2 is a flow chart of a model training method provided in an embodiment of the present application; the model training method 200 includes the following steps:

[0087] 201. A model training device obtains a sample set, wherein the sample set includes a plurality of first images, and an inspection item label and an inspection result of each of the plurality of first images.

[0088] Specifically, the first image is an image related to the inspection site. The inspection item label is used to indicate the category of the inspection item in the first image. The inspection item category is the specific type of inspection item, for example, the inspection item can be a base station main device, an outdoor antenna, or an air conditioner. The inspection result is used to indicate the inspection result corresponding to the inspection item in the first image. For example, the inspection result can be used to indicate the operating status of the inspection item, which can include normal status or abnormal status.

[0089] 202. The model training device performs model training based on the sample set and the first inspection model to obtain a second inspection model. The first inspection model is used for inspection processing of at least one first inspection item in an inspection scenario. The inspection scenario is determined based on multiple first images. The at least one first inspection item intersects with the inspection item corresponding to the sample set.

[0090] Specifically, after determining the inspection scene based on multiple first pictures, model training is performed based on the sample set and the first inspection model corresponding to the inspection scene to obtain the second inspection model.

[0091] The first inspection model and the second inspection model are used for the inspection process of the inspection site. Further, the first inspection model and the second inspection model correspond to the inspection process of the inspection site in the inspection scenario.

[0092] Exemplarily, the inspection site may include at least one inspection scene. The embodiment of the present application matches the inspection scene based on the first image to determine the corresponding first inspection model for model training. It can conduct targeted training of inspection models for different inspection scenes, which helps to improve the accuracy of the inspection model.

[0093] In the embodiment of the present application, not only can the inspection scenario type be limited, but prior knowledge (i.e., sample set) can be applied to assist model training. Moreover, since the first inspection model is a preliminarily trained inspection model, the training time required to obtain the second inspection model based on the first inspection model is relatively short, which can improve the efficiency of inspection model training.

[0094] In one possible embodiment, in step 201, the model training device may obtain an inspection result image from the inspection device as a first image. The model training device may periodically or quantitatively collect multiple first images as a sample set. If the inspection result image does not include inspection item labels and inspection results, the model training device may annotate the inspection result image to obtain the first image. Exemplarily, the model training device may construct the sample set using either intelligent or manual annotation.

[0095] When the model training device uses the intelligent labeling method to label multiple inspection result images, it specifically includes the following steps S1 to S2:

[0096] S1. The model training device uses the basic detection model to infer the inspection result image to be labeled, and obtains the inferred label of the inspection result image. The inferred label includes the inspection item label and the inspection result of the inspection result image. The inspection items targeted by the above basic detection model are the inspection items included in the historical inspection scene.

[0097] S2. Manually review and confirm the above-mentioned inference labels. If the above-mentioned inference labels are manually confirmed to be qualified (i.e., accurate), the above-mentioned inference labels are used as the annotation data of the above-mentioned inspection result images, and the inference is continued on the remaining inspection result images. If the above-mentioned inference labels are manually confirmed to be unqualified (i.e., inaccurate), the above-mentioned inference labels are manually corrected. The model training device optimizes the basic detection model based on the inspection result images after label correction, and then uses the optimized basic detection model to reason on the remaining inspection result images. The above process is repeated until all inspection result images are labeled, and the labeled inspection result images are used to construct a sample set.

[0098] When manual labeling is used, the platform labeling tools are used to label the dataset to be labeled and complete the construction of the sample set.

[0099] For example, when the inspection result picture includes inspection item labels and inspection results, if there is no need to add new inspection items, the inspection result picture can be directly used as training data for the model. If new inspection items need to be added, the inspection result picture needs to be annotated manually.

[0100] In a possible implementation, the model training method 200 further includes the following steps S3 to S4:

[0101] S3. The model training device determines the inspection scenarios corresponding to some or all of the plurality of first images. Assuming the plurality of first images is N first images, the portion of the first images may be N*a% first images. The specific value of a can be set based on actual conditions and is not limited.

[0102] S4. The model training device uses the inspection scene with the largest number of images as the inspection scene corresponding to the sample set.

[0103] Specifically, when determining the inspection scenes corresponding to all of the first pictures among the multiple first pictures, the inspection scene with the largest number of pictures among the multiple inspection scenes is used as the inspection scene corresponding to the sample set. Furthermore, when determining the inspection scenes corresponding to N*a% of the first pictures among the multiple first pictures, the inspection scene with the largest number of pictures among the N*a% inspection scenes is used as the inspection scene corresponding to the sample set.

[0104] In the embodiment of the present application, a voting mechanism is used to determine the inspection scenario corresponding to the sample set to ensure the accuracy of the inspection scenario. Further, the first inspection model corresponding to the sample set at this time can be determined based on the inspection scenario.

[0105] For example, based on the inspection scenario, training configuration parameters corresponding to the first inspection model can also be determined. The training configuration parameters indicate configuration information related to the training of the first inspection model, such as the number of training rounds for the first inspection model, the number of images processed by the first inspection model during each training session, the image scaling factor, and the sample augmentation technology used by the first inspection model. Accordingly, in step 202, when training the model, model training can be performed based on the training configuration parameters determined above.

[0106] In one possible embodiment, the model training device uses a scene classification model to classify the first image in the sample set, and then votes based on the classification results of the first image to obtain the inspection scene of the sample set, thereby determining the first inspection model (and corresponding training configuration parameters) for model training. Assuming that the total number of first images in the sample set is N, the specific steps include the following steps S5 to S7:

[0107] S5. Randomly extract a% of the first images from the sample set and sequentially input the extracted first images into the scene classification model to infer the inspection scene corresponding to each first image. In this embodiment, the specific value of a can be set according to actual conditions and is not particularly limited. For example, a can be 80.

[0108] S6. Count the number of inspection scenes in the inference results and select the inspection scene with the largest number as the inspection scene corresponding to the sample set.

[0109] S7. Determine a corresponding first inspection model (and corresponding training configuration parameters) based on the inspection scenario of the sample set.

[0110] In the embodiment of the present application, a voting mechanism is used to determine the inspection scenario corresponding to the sample set, which can improve the accuracy of the inspection scenario.

[0111] In a possible implementation, the model training method 200 further includes the following steps S8 to S11:

[0112] S8. The model training device determines k first images corresponding to each inspection item label in the sample set, where k is greater than or equal to 1.

[0113] Specifically, the specific value of k can be set according to actual conditions and is not particularly limited. For example, k is 3.

[0114] S9. The model training device crops k second images of the area where the inspection items are located from the k first images of each inspection item label.

[0115] Specifically, the second image of the inspection item label is an image of the inspection item corresponding to the inspection item label, and the second image can be the first image itself or a local area image of the first image. When the second image is the first image itself, this step can be skipped.

[0116] When the second image is a partial image of the first image, the size of the second image is smaller than that of the first image. For example, the second image can be obtained by cropping the first image using a first marker in the first image, where the first marker is used to indicate the specific location of the inspection item in the first image. For example, the first marker can be a frame that can be used to frame the inspection item in the first image.

[0117] S10. The model training device uses a feature extraction network to perform feature extraction on the k second images of each inspection item label to obtain k second features of each inspection item label.

[0118] S11. The model training device obtains the first feature of each inspection item label based on the k second features of each inspection item label.

[0119] Specifically, the specific method of fusing k second features to obtain the first feature can be to directly add the k second features to obtain the first feature, or to multiply each second feature by a coefficient and then add them together. The specific size of the above coefficient can be set according to actual conditions.

[0120] Exemplarily, when the number of first images corresponding to the inspection item label is less than k, the same processing as above is performed on all first images corresponding to the inspection item label to obtain the first feature.

[0121] In the embodiment of the present application, it is assumed that the first features corresponding to all inspection item labels constitute a first feature set F s , where F s ={f s1 ,f s2 …f sl}, where l is the number of categories of inspection item labels in the sample set. Using the baseline feature of the first inspection item and each first feature, l first similarities are calculated. The inspection item label corresponding to the first feature with the highest first similarity is then assigned to the first inspection item. Performing this process on each inspection item label yields the first feature of each inspection item label.

[0122] In one possible implementation, referring to FIG. 3A , FIG. 3A is a schematic diagram of a specific flow of step 202 provided in an embodiment of the present application; the above-mentioned step 202 specifically includes the following steps 221 to 223:

[0123] 221. The model training device determines a correspondence between the inspection item label and the at least one first inspection item based on the sample set and the reference inspection picture of each first inspection item in the at least one first inspection item.

[0124] Specifically, since the inspection item label and the name of the first inspection item may be different, the baseline inspection image of the first inspection item and the first image in the sample set can be used to determine whether there is a correspondence between the inspection item label and the first inspection item. Determining whether the inspection item label and the first inspection item are similar (i.e., corresponding) based on the image makes the judgment result more accurate and reliable.

[0125] Exemplarily, the above step 221 specifically includes the following steps S12 to S14:

[0126] S12: The model training device uses a feature extraction network to extract features from the reference inspection image of the second inspection item to obtain a first reference feature. At least one first inspection item includes the second inspection item.

[0127] Specifically, each first inspection item has at least one reference inspection picture. The reference inspection picture of the first inspection item is a picture including the first inspection item, which can be understood as a standard picture including the first inspection item. The reference inspection picture can be pre-stored.

[0128] When the first inspection item has more than two reference inspection pictures, feature extraction and feature fusion may be performed on each reference inspection picture to obtain a first reference feature.

[0129] S13. The model training device determines a first similarity between each inspection item label and a second inspection item based on the first benchmark feature and the first feature of each inspection item label in the sample set. The first feature of the inspection item label is determined based on the image area where the inspection item is located in the first image.

[0130] Specifically, the model training device can process the image area of ​​the inspection item in at least one first image corresponding to the inspection item label to obtain the first feature corresponding to the inspection item label. In this embodiment, the cosine distance or Euclidean distance between the first reference feature and the first feature is used as the first similarity.

[0131] S14. The model training device determines that the above correspondence relationship includes that the inspection item label with the highest first similarity among the inspection item labels corresponds to the second inspection item.

[0132] In an embodiment of the present application, the first similarity between the first reference feature and the first feature of the inspection item label is determined, and the inspection item label with the highest first similarity among the inspection item labels is matched with the second inspection item, thereby determining the correspondence between the inspection item label and the first inspection item, so as to facilitate the determination of the training strategy of the first inspection model. In addition, the same method as the above-mentioned second inspection item is adopted for each first inspection item in at least one first inspection item in the embodiment of the present application to complete the matching (i.e., correspondence) of each first inspection item with the inspection item label. In an embodiment of the present application, feature extraction is performed on the image to obtain the representative feature, and the first similarity can be obtained more accurately, and then based on the first similarity, it can be accurately determined whether the inspection item label and the second inspection item correspond.

[0133] For example, the model training device determines that an inspection item label having a first similarity greater than a similarity threshold corresponds to a second inspection item. The specific value of the similarity threshold can be set based on actual conditions and is not particularly limited. For example, the similarity threshold is 80%, 95%, or 98%.

[0134] 222. The model training device determines a training strategy for the first inspection model based on the corresponding relationship.

[0135] Exemplarily, the above step 222 specifically includes the following steps S15 to S18:

[0136] S15. Determine at least one third inspection item of the first inspection model based on the corresponding relationship; the at least one third inspection item is an inspection item that requires training of the first inspection model.

[0137] Specifically, after the model training device automatically completes the correspondence between the first inspection item and the inspection item label using steps S12 to S14 above, it is necessary to manually confirm whether the correspondence between the first inspection item and the inspection item label is correct, and manually determine the third inspection item based on the correspondence between the first inspection item and the inspection item label; the third inspection item can be selected from the inspection items corresponding to the first inspection item and the inspection item label. Based on the above correspondence, the inspection personnel can select a certain first inspection item and the inspection item corresponding to the corresponding inspection item label as a single inspection item.

[0138] For example, referring to FIG. 3B , FIG. 3B is a schematic diagram of determining a third inspection item provided by an embodiment of the present application. In FIG. 3B , the "Outdoor Antenna Item Label Name," the "Base Station Main Equipment Item Label Name," and the "Dummy Panel Missing Item Label Name" are the first inspection items of the first inspection model. The model training device uses steps S12 to S14 above to fill the inspection item label corresponding to the first inspection item in the sample set into the box corresponding to the first inspection item, indicating that the first inspection item and the inspection item label have a corresponding relationship. Inspection item labels that cannot be matched to the first inspection item in the sample set are filled into the box corresponding to "Add Inspection Item."

[0139] The inspection personnel can select the third inspection item from the inspection items corresponding to the first inspection item and the inspection item label according to the actual situation. For example, the inspection personnel can perform the inspection item selection operation after pulling down the box in Figure 3B to determine the third inspection item. When the correspondence between the first inspection item and the inspection item label is correct, the inspection personnel can select the first inspection item as the third inspection item. Alternatively, when the correspondence between the first inspection item and the inspection item label is wrong, the inspection personnel may not use the first inspection item as the third inspection item. Alternatively, when the correspondence between the first inspection item and a certain inspection item label is wrong, the inspection personnel can cancel the correspondence between the first inspection item and the inspection item label, and then use the first inspection item as the third inspection item. As for the inspection item label that cannot correspond to the first inspection item (including the inspection item label whose correspondence has been canceled), the inspection personnel can determine the inspection items corresponding to several inspection item labels as the third inspection item according to the actual situation.

[0140] Further exemplarily, referring to FIG3B , in an embodiment of the present application, inspection personnel may also determine whether to generate a terminal-side model and whether to start incremental training.

[0141] S16. When at least one third inspection item corresponds one-to-one to at least one first inspection item, the training strategy is to load the model parameters of the first inspection model and then perform model training.

[0142] Specifically, the first inspection model is a preliminarily trained model, and therefore has corresponding model parameters. When all third inspection items correspond one-to-one with all first inspection items, the training strategy for the first inspection model is to load all model parameters and then perform model training using the sample set, thereby adjusting all model parameters of the first inspection model.

[0143] S17. When at least one third inspection item corresponds one-to-one to part of the inspection items of at least one first inspection item, the training strategy is to load the model parameters of the first inspection model, keep the parameters of the backbone network in the first inspection model unchanged, and perform model training on the detection head in the first inspection model.

[0144] Specifically, the first inspection model includes a backbone network and a detection head. The backbone network is used to extract features from the input image to obtain representation features, while the detection head is used to classify and locate targets in the input image using the above representation features.

[0145] When all the third inspection items and some of the first inspection items correspond one to one (the set of third inspection items is a true subset of the set of first inspection items), the training strategy of the first inspection model is to load all the model parameters and then use the sample set to train the model, wherein the parameters of the backbone network in the first inspection model are kept unchanged, that is, the model parameters of the backbone network in the first inspection model cannot be adjusted, but the model parameters of the detection head in the first inspection model can be adjusted.

[0146] S18. When some inspection items of at least one third inspection item correspond one-to-one to some inspection items of at least one first inspection item, the training strategy is to load the parameters of the backbone network in the first inspection model, and randomly initialize the parameters of the detection head in the first inspection model before performing model training.

[0147] Specifically, when some of the third inspection items correspond one-to-one to some of the first inspection items, the training strategy of the first inspection model is to load the model parameters of the backbone network in the first inspection model, and after randomly initializing the parameters of the detection head, use the sample set for model training, and all the model parameters of the first inspection model can be adjusted.

[0148] 223. The model training device trains the first training model based on the sample set and the training strategy to obtain a second inspection model.

[0149] In an embodiment of the present application, after determining the correspondence between the first inspection item and the inspection item label, at least one third inspection item of the first inspection model can be determined based on the correspondence, and the third inspection item can be determined from the inspection items corresponding to the first inspection item and the inspection item label. A training strategy for the first inspection model is then determined based on the correspondence between the third inspection item and the first inspection item. Different model training strategies can be determined to suit the specific circumstances of the third inspection item, thereby improving the speed of model training and the accuracy of the second inspection model obtained after training.

[0150] Specifically, the model training device performs model training based on the sample set and the determined training strategy. When the training meets the stopping condition, the training automatically stops and the second inspection model generated by the training is sent to the inspection device for updating. For example, in this embodiment, the stopping condition is set to when the cumulative number of training rounds reaches a set total number of rounds.

[0151] In one possible implementation, referring to Figure 1A, the inspection device performs inspection operations based on the second inspection model trained by the model training device. The inspection device can call different second inspection models to perform inspection operations based on different inspection scenarios.

[0152] For example, for a computer room scenario, the second inspection model corresponding to the computer room can be called to perform inspection processing. For an outdoor antenna scenario, the second inspection model corresponding to the outdoor antenna can be called to perform inspection processing. As another example, the inspection equipment can perform scene recognition based on a picture of the inspection site to determine the current inspection scene, and then determine the second inspection model based on the scene. Alternatively, the inspection equipment can determine the inspection scene based on the inspection items in the inspection work order, and then determine the second inspection model based on the scene. Alternatively, the inspection personnel can manually select the corresponding second inspection model on the inspection equipment based on the inspection scene to perform inspection operations.

[0153] In one possible implementation, when the number of first images corresponding to each inspection item in some or all of the at least one third inspection item is less than or equal to a quantity threshold, the first inspection model is a small sample inspection model. The specific value of the quantity threshold can be set based on actual circumstances and is not specifically limited. Similarly, the specific number of some inspection items can be set based on actual circumstances and is not specifically limited.

[0154] A small sample inspection model refers to an inspection model that can complete model training using a smaller number of samples. When the number of sample images for all or part of the third inspection items is small, the first inspection model can use the small sample inspection model to successfully complete model training and obtain the second inspection model. The embodiment of the present application supports the selection of whether to call the small sample inspection model for model training based on the amount of sample data corresponding to the inspection item, which can reduce the demand for training data volume, quickly and efficiently iterate the model, and improve the detection accuracy of the inspection model.

[0155] For the first inspection models of different inspection scenarios, corresponding small sample inspection models can be respectively set to adapt to different inspection scenarios.

[0156] Exemplarily, the small sample inspection model includes a prototype network module. Further exemplarily, the small sample inspection model also includes a prototype enhancement module obtained by performing an affine transformation on the prototype network module.

[0157] For example, the first inspection model includes a backbone network and a detection head, while the small sample inspection model corresponding to the first inspection model includes, in sequence, a backbone network, a prototype network module, a network layer (the specific structure of the network layer is not particularly limited), a prototype enhancement module, and a detection head. The prototype network module processes the output features of the backbone network to obtain more recognizable features. The processed features are then transformed through the network layer and then processed by the prototype enhancement module to obtain more recognizable image features. These features are then sent to the detection head for output of the final detection results.

[0158] The above describes in detail the method of the embodiment of the present application. The following describes the device provided by the embodiment of the present application.

[0159] The present application also provides a device, which includes a unit or module for executing the method described in any of the above embodiments.

[0160] FIG4 is a schematic diagram of the structure of a possible model training device provided in an embodiment of the present application. The model training device shown in FIG4 can be used to implement the functions of the above-mentioned model training method embodiment, and thus can also achieve the beneficial effects possessed by the above-mentioned model training method embodiment. In an embodiment of the present application, the model training device can be an electronic device, or a module (such as a chip) used in an electronic device.

[0161] As shown in Figure 4, the model training device 400 includes an acquisition module 401 and a training module 402. The model training device 400 is used to implement the functions of the model training method embodiment shown in Figure 2 above. Alternatively, the model training device 400 may include a module for implementing any function or operation in the model training method embodiment shown in Figure 2 above, and the module may be implemented in whole or in part by software, hardware, firmware, or any combination thereof.

[0162] When the model training device 400 is used to implement the functions of the model training method embodiment shown in Figure 2, the acquisition module 401 is used to obtain a sample set. The above-mentioned sample set includes multiple first pictures, and the inspection item label and inspection result of each first picture in the multiple first pictures. The training module 402 is used to perform model training based on the sample set and the first inspection model to obtain a second inspection model; the first inspection model is used for inspection processing of at least one first inspection item in the inspection scenario. The above-mentioned inspection scenario is determined based on multiple first pictures. The above-mentioned at least one first inspection item has an intersection with the inspection item corresponding to the sample set.

[0163] In one possible implementation, the training module 402 is specifically configured to: determine a correspondence between an inspection item label and the at least one first inspection item based on the sample set and a baseline inspection image for each of the at least one first inspection item; determine a training strategy for the first inspection model based on the correspondence; and train the first inspection model based on the sample set and the training strategy to obtain a second inspection model.

[0164] In one possible embodiment, the training module 402 is specifically used to determine the correspondence between the inspection item label and at least one first inspection item based on the sample set and the benchmark image of at least one first inspection item: perform feature extraction on the benchmark inspection image of the second inspection item using a feature extraction network to obtain a first benchmark feature. At least one first inspection item includes the second inspection item. Based on the first benchmark feature and the first feature of each inspection item label in the sample set, determine the first similarity between each inspection item label and the second inspection item. The first feature of the inspection item label is determined based on the image area where the inspection item is located in the first image. Determining the above correspondence includes that the inspection item label with the highest first similarity among the inspection item labels corresponds to the second inspection item. Wherein, the benchmark inspection image is a picture including the above-mentioned second inspection item.

[0165] In one possible embodiment, the above-mentioned model training device 400 also includes a determination module 403, an extraction module 404 and a fusion module 405. The determination module 403 is used to determine the k first images corresponding to each inspection item label in the sample set, where k is greater than or equal to 1. The acquisition module 401 is also used to crop the second images of the area where the k inspection items are located from the k first images of each inspection item label. The extraction module 404 is used to use a feature extraction network to perform feature extraction on the k second images of each inspection item label to obtain k second features of each inspection item label. The fusion module 405 is used to obtain the first feature of each inspection item label based on the fusion of the k second features of each inspection item label.

[0166] In one possible embodiment, the above-mentioned training module 402, in terms of determining the training strategy of the first inspection model based on the corresponding relationship, is specifically used to: determine at least one third inspection item of the first inspection model based on the corresponding relationship; at least one third inspection item is an inspection item that requires the first inspection model to be trained. When at least one third inspection item corresponds one-to-one with at least one first inspection item, the training strategy is to load the model parameters of the first inspection model and then perform model training. When at least one third inspection item corresponds one-to-one with part of the inspection items of at least one first inspection item, the training strategy is to load the model parameters of the first inspection model, keep the parameters of the backbone network in the first inspection model unchanged, and perform model training on the detection head in the first inspection model. When part of the inspection items of at least one third inspection item corresponds one-to-one with part of the inspection items of at least one first inspection item, the training strategy is to load the parameters of the backbone network in the first inspection model, and randomly initialize the parameters of the detection head in the first inspection model and then perform model training.

[0167] In one possible implementation, when the number of first images corresponding to each inspection item in some or all of the at least one third inspection item is less than or equal to a quantity threshold, the first inspection model is a small sample inspection model. A small sample inspection model refers to an inspection model that can complete model training using a relatively small number of samples.

[0168] Exemplarily, the small sample inspection model includes a prototype network module.

[0169] In a possible implementation, the determination module 403 is further configured to determine the inspection scenes corresponding to some or all of the first pictures among the plurality of first pictures, and to use the inspection scene with the largest number of pictures as the inspection scene corresponding to the sample set.

[0170] For the introduction of the above modules, please refer to the description of the above embodiments, which will not be repeated here.

[0171] Referring to Figure 5, Figure 5 is a schematic diagram of the structure of another model training device provided in an embodiment of the present application. The present application also provides a model training device, wherein the model training device 500 includes a memory 501, a processor 502, a communication interface 504, and a bus 503. The memory 501, the processor 502, and the communication interface 504 are connected to each other via the bus 503.

[0172] The memory 501 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 501 may store a program. When the program stored in the memory 501 is executed by the processor 502, the processor 502 and the communication interface 504 are used to perform the various steps of the model training method of any embodiment of the present application.

[0173] The processor 502 can adopt a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU) or one or more integrated circuits to execute relevant programs to implement the functions required to be performed by the units in the model training device of any embodiment of the present application, or to execute the model training method of any embodiment of the present application.

[0174] The processor 502 may also be an integrated circuit chip with signal processing capabilities. During implementation, the various steps of the model training method of any embodiment of the present application may be completed by hardware integrated logic circuits or software instructions in the processor 502. The above-mentioned processor 502 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application may be implemented or executed. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the model training method in conjunction with any embodiment of the present application may be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 501, and the processor 502 reads the information in the memory 501 and combines its hardware to complete the functions required to be performed by the units included in the model training device of any embodiment of the present application, or executes the model training method of any embodiment of the present application.

[0175] The communication interface 504 uses a transceiver such as, but not limited to, a transceiver to implement communication between the model training device 500 and other devices or a communication network. For example, the first image can be obtained through the communication interface 504.

[0176] The bus 503 may include a path for transmitting information between the various components of the model training device 500 (e.g., memory 501, processor 502, communication interface 504). In the several embodiments provided in this application, it should be understood that the disclosed apparatus and method can be implemented in other ways. For example, the apparatus embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices or units, which may be electrical, mechanical or other forms.

[0177] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0178] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0179] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a read-only memory (ROM), a random access memory (RAM), or a magnetic medium, such as a floppy disk, a hard disk, a tape, a magnetic disk, or an optical medium, such as a digital versatile disc (DVD), or a semiconductor medium, such as a solid state drive (SSD).

[0180] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A model training method, characterized in that: include: Acquire a sample set, the sample set comprising a plurality of first images, and an inspection item label and an inspection result of the inspection item of each first image in the plurality of first images; Performing model training based on the sample set and the first inspection model to obtain a second inspection model; The first inspection model is used for inspection processing of at least one first inspection item in an inspection scenario, where the inspection scenario is determined based on the multiple first images; and the at least one first inspection item has an intersection with the inspection item corresponding to the sample set.

2. The method according to claim 1, characterized in that The performing model training based on the sample set and the first inspection model to obtain a second inspection model includes: Determine a correspondence between the inspection item label and the at least one first inspection item based on the sample set and a reference inspection picture of each first inspection item in the at least one first inspection item; Determining a training strategy for the first inspection model based on the corresponding relationship; The first inspection model is trained based on the sample set and the training strategy to obtain the second inspection model.

3. The method according to claim 2, characterized in that The determining, based on the sample set and the reference picture of the at least one first inspection item, a corresponding relationship between the inspection item label and the at least one first inspection item includes: Using a feature extraction network to extract features from a reference inspection picture of a second inspection item to obtain a first reference feature; the at least one first inspection item includes the second inspection item; Determine a first similarity between each inspection item label and the second inspection item based on the first reference feature and a first feature of each inspection item label in the sample set, wherein the first feature of the inspection item label is determined based on an image region where the inspection item is located in the first image; Determining the corresponding relationship includes that the inspection item label with the highest first similarity among the inspection item labels corresponds to the second inspection item.

4. The method according to claim 3, characterized in that The method further comprises: Determine k first pictures corresponding to each inspection item label in the sample set, where k is greater than or equal to 1; Cut out second images of areas where k inspection items are located from the k first images of each inspection item label; Using the feature extraction network to perform feature extraction on the k second images of each inspection item label to obtain k second features of each inspection item label; The first feature of each inspection item label is obtained by fusing the k second features of each inspection item label.

5. The method according to any one of claims 2 to 4, characterized in that: The determining of the training strategy of the first inspection model based on the corresponding relationship includes: Determine at least one third inspection item of the first inspection model based on the corresponding relationship; the at least one third inspection item is an inspection item that requires training of the first inspection model; When the at least one third inspection item corresponds one-to-one to the at least one first inspection item, the training strategy is to perform model training after loading the model parameters of the first inspection model; When the at least one third inspection item corresponds one-to-one to some inspection items of the at least one first inspection item, the training strategy is to load the model parameters of the first inspection model, keep the parameters of the backbone network in the first inspection model unchanged, and perform model training on the detection head in the first inspection model; When some inspection items of at least one third inspection item correspond one-to-one to some inspection items of at least one first inspection item, the training strategy is to load the parameters of the backbone network in the first inspection model, and randomly initialize the parameters of the detection head in the first inspection model before performing model training.

6. The method according to claim 5, characterized in that When the number of first images corresponding to each inspection item in some or all of the at least one third inspection item is less than or equal to a quantity threshold, the first inspection model is a small sample inspection model.

7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: Determining inspection scenes corresponding to some or all of the first images among the plurality of first images; The inspection scene with the largest number of pictures is taken as the inspection scene corresponding to the sample set.

8. A model training device, characterized in that: The device includes a unit or module for executing the model training method described in any one of claims 1 to 7.

9. A model training device, characterized in that: It includes a processor and a memory, wherein the processor and the memory are connected, wherein the memory is used to store program code, and the processor is used to call the program code to execute the model training method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the model training method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Inspection method and device

    CN114783188A

  • Anomaly detection method and anomaly detection system

    CN116861346A

  • Method and device for training prediction model for new scenario

    WO2020024716A1