Method and device for realizing image detection based on computing power of intelligent computing center
By acquiring and converting image object area information on the intelligent computing center and training the object detection model, the low detection accuracy problem caused by the lack of labeled data is solved, and efficient image detection is achieved.
Patent Information
- Application Number
- CN202510331435.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-17
AI Technical Summary
In the scenario where labeled data is lacking, the detection accuracy is very low when the intelligent computing center dispatches computing resources for image detection.
By obtaining object area information in the image, converting it into a format that the object detection model can recognize, and using this information to train the object detection model, thereby improving the accuracy of image detection.
By using the trained object detection model for image detection, the detection accuracy of images can be greatly improved, the workload of manual labeling and the labeling efficiency can be improved.
Smart Images

Figure CN120163971A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of intelligent computing centers, intelligent computing centers and computing power infrastructure, and particularly relates to a method and device for realizing image detection based on the computing power of an intelligent computing center. Background Art
[0002] With the rapid development of artificial intelligence technology, "intelligent computing centers" and "intelligent computing centers" have emerged as the times require.
[0003] An "intelligent computing center" refers to a facility that uses large-scale heterogeneous computing power resources, including general computing power and intelligent computing power, and mainly provides the required computing power, data and algorithms for artificial intelligence applications (such as scenarios of artificial intelligence deep learning model development, model training and model inference, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0004] The "intelligent computing center" includes, but is not limited to, the "intelligent computing center".
[0005] An "intelligent computing center", that is, an artificial intelligence computing center, is a type of computing power infrastructure that provides computing power services, data services and algorithm services required for artificial intelligence applications based on artificial intelligence theory and adopting an artificial intelligence computing architecture.
[0006] "Computing power" is the core of "intelligent computing centers" and "intelligent computing centers", which is the ability of computer devices or computing / data centers to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of target results by processing information data, and a new type of productive force integrating information computing power, network carrying capacity and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0007] Currently, usually, the training of an image detection model requires a large amount of labeled data. However, in actual applications, there is often a lack of labeled data, especially in some scenarios such as new categories or new fields. In this case, due to the lack of labeled data for training, when the intelligent computing center schedules computing power resources for image detection, the detection accuracy is very low. Summary of the Invention
[0008] The present invention provides a method and device for realizing image detection based on the computing power of an intelligent computing center, which is used to solve the problem that when the intelligent computing center schedules computing power resources for image detection in a scenario lacking labeled data, the detection accuracy is very low.
[0009] To solve the above technical problems, the present invention is implemented as follows:
[0010] In a first aspect, the present invention provides a method for implementing image detection based on the computing power of an intelligent computing center, including:
[0011] Step S1: Obtain a first image;
[0012] Step S2: Identify a first region in the first image where a first object indicated by a first keyword input is located;
[0013] Step S3: Obtain information of the first region and convert the information of the first region into first region information in a target format, where the target format is a format recognizable by a target detection model;
[0014] Step S4: Establish a correspondence between first category information to which the first object belongs and the first region information;
[0015] Step S5: Use the first category information, the first region information, and the correspondence to train the target detection model, where the target detection model is used to implement image detection based on the computing power of the intelligent computing center.
[0016] Optionally, after step S5, the method further includes:
[0017] Step S6: Input a second image to be detected and a second keyword into the trained target detection model, where the second keyword is used to indicate the first object;
[0018] Step S7: Use the trained target detection model to detect the category of the first object indicated by the second keyword and second region information of the first object in the second image, and output the category information of the first object and the second region information.
[0019] Optionally, step S7 includes:
[0020] Step S71: Use the trained target detection model to determine the category of at least one first object indicated by the second keyword and at least one second region information of the at least one first object in the second image, and a confidence level corresponding to each second region information;
[0021] Step S72: Output the category information of the first object with a confidence level greater than a preset threshold and the corresponding second region information.
[0022] Optionally, after step S5, the method further includes:
[0023] Step S8: Generate an identifier for the target detection model, and establish a mapping relationship among the identifier, the storage path of the target detection model, and a keyword list. The keyword list includes keywords input into the target detection model, and the second keyword is included in the keyword list;
[0024] The step S6 includes:
[0025] Step S61: When receiving the input second keyword, determine the target detection model corresponding to the second keyword according to the mapping relationship;
[0026] Step S62: Input the second image to be detected and the second keyword into the trained target detection model.
[0027] Optionally, the step S1 includes at least one of the following:
[0028] Step S11: Determine the first image associated with the third keyword in the pre-stored images according to the input third keyword and the preset association relationship between the image name and the keyword stored in advance;
[0029] Step S12: Retrieve the first image associated with the fourth keyword in the server according to the input fourth keyword;
[0030] Step S13: Generate the first image associated with the fifth keyword according to the input fifth keyword.
[0031] Optionally, the step S3 includes:
[0032] Step S31: Obtain the pixel mask or the regional boundary contour of the first region, and determine the minimum bounding rectangle of the region corresponding to the pixel mask or the regional boundary contour to obtain the information of the minimum bounding rectangle;
[0033] Step S32: Convert the format of the information of the minimum bounding rectangle into the target format, and perform normalization processing on the information of the minimum bounding rectangle. After the normalization processing, the information of the minimum bounding rectangle is within the target interval, and the first region information in the target format includes the information of the minimum bounding rectangle after the normalization processing.
[0034] In a second aspect, the present invention provides an apparatus for implementing image detection based on the computing power of an intelligent computing center, including:
[0035] An acquisition module, configured to acquire a first image;
[0036] An identification module, configured to identify the first region where the first object indicated by the input first keyword is located in the first image;
[0037] A conversion module, configured to obtain information of the first region and convert the information of the first region into first region information in a target format, where the target format is a format recognizable by a target detection model;
[0038] An establishment module, configured to establish a correspondence relationship between the first category information to which the first object belongs and the first region information;
[0039] A training module, configured to train the target detection model by using the first category information, the first region information, and the correspondence relationship, where the target detection model is used to implement image detection based on the computing power of an intelligent computing center.
[0040] In a third aspect, the present invention provides an electronic device, including: a processor, a memory, and a program stored on the memory and executable on the processor, where when the program is executed by the processor, the steps of the method for implementing image detection based on the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0041] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, where when the computer program is executed by a processor, the steps of the method for implementing image detection based on the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0042] In a fifth aspect, the present invention provides a computer program product, including computer instructions, where when the computer instructions are executed by a processor, the steps of the method for implementing image detection based on the computing power of an intelligent computing center as described in the first aspect above are implemented.
[0043] In the present invention, by using the computing power of an intelligent computing center, the first object in an image is recognized and converted into first region information recognizable by a target detection model, and the target detection model is trained by using the first region information and the category information of the first object, so as to obtain a target detection model for implementing image detection based on the computing power of an intelligent computing center. By training the target detection model by using the category information and region information of the first object and performing image detection by using the trained target detection model based on the computing power of an intelligent computing center, the detection accuracy of the image can be greatly improved. Description of the Drawings
[0044] By reading the following detailed description of the preferred embodiments, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. Moreover, throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:
[0045] Figure 1 It is a schematic flow diagram of the method for implementing image detection based on the computing power of the intelligent computing center of the present invention;
[0046] Figure 2 It is a schematic structural diagram of the device for implementing image detection based on the computing power of the intelligent computing center of the present invention;
[0047] Figure 3 It is a schematic structural diagram of the electronic device of the present invention. Specific embodiments
[0048] Next, the technical solutions in the present invention will be clearly and completely described in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0049] First, the technical terms related to the present invention will be briefly described below.
[0050] The "computing power" referred to in the present invention means: It is the ability of a computer device or a computing / data center to process information, the ability of computer hardware and software to cooperate to jointly execute a certain computing requirement, the computing ability to achieve the output of the target result by processing information data, a new type of productive force integrating information computing power, network carrying capacity, and data storage capacity, and mainly provides services to society through computing power infrastructure.
[0051] The "computational power" (Computational Power, CP) referred to in the present invention means: It is the ability of a data center server to process data and achieve the output of the result, a comprehensive index to measure the computing ability of a data center, including general computing ability, supercomputing ability, and intelligent computing ability. The commonly used measurement unit is the number of floating-point operations per second (FLOPS, 1EFLOPS = 10^18 FLOPS), and the larger the value, the stronger the comprehensive computing ability. It is estimated that 1EFLOPS is approximately the computing power output of 5 Tianhe 2A or 500,000 mainstream server CPUs or 2 million mainstream laptops. The calculation formula is: CP = CP_general + CP_intelligent + CP_super.
[0052] The "carrying capacity" (Network Power, NP) referred to in the present invention means: It is the manifestation of the data transmission ability of the computing power facility, a comprehensive ability including network architecture, network bandwidth, transmission delay, intelligent management and scheduling, etc., involving network transmission inside and between data centers, and is a comprehensive index to measure the network transmission scheduling ability.
[0053] The "Storage Power (SP)" described in the present invention refers to the comprehensive ability of a data center in four aspects: data storage capacity, performance, security and reliability, and green and low-carbon. It is a comprehensive indicator for measuring the data storage capacity of a data center, including external storage devices such as storage arrays and built-in storage devices of servers. The commonly used measurement unit for storage capacity is exabyte (EB, 1EB = 2^60 bytes), the commonly used measurement unit for performance is the number of read and write operations per second per unit capacity (IOPS / TB, Input / Output Operations Per Second / TB), and the disaster recovery ratio is an important manifestation of security and reliability.
[0054] The "computing power infrastructure" described in the present invention refers to a new type of information infrastructure that integrates information computing power, network carrying capacity, and data storage power, and can realize centralized computing, storage, transmission, and application of information.
[0055] The "new type of information infrastructure" described in the present invention mainly includes network infrastructures such as 5G networks, fiber broadband networks, backbone networks, international communication networks, and satellite Internet, computing power infrastructures such as data centers, general computing power centers, intelligent computing centers, and supercomputing centers, and new technology facilities such as artificial intelligence, blockchain, and quantum computing.
[0056] The "computing power" described in the present invention includes: general computing power, intelligent computing power, and super computing power.
[0057] The "general computing power" described in the present invention refers to the computing ability provided by servers based on CPU (Central Processing Unit) chips, which is used to support basic general computing such as cloud computing and edge computing.
[0058] The "intelligent computing power" described in the present invention refers to a computing platform for large-scale deployment of dedicated chips such as GPU (Graphics Processing Unit), FPGA (Field Programmable Gate Array), and ASIC (Application Specific Integrated Circuit) for various artificial intelligence innovation applications, such as Natural Language Processing (NLP), machine vision, etc.
[0059] The "super computing power" described in the present invention refers to: mainly the computing power provided by high-performance computing clusters such as supercomputers. It utilizes the centralized computing resources of multiple computer systems working in parallel and processes extremely complex or data-intensive problems through a dedicated operating system. It is mainly used for computing in cutting-edge scientific fields, such as planetary simulation, drug molecule design, gene analysis, etc.
[0060] The "intelligent computing center" described in the present invention refers to: a facility that provides the required computing power, data, and algorithms mainly for artificial intelligence applications (such as scenarios like artificial intelligence deep learning model development, model training, and model inference) by using large-scale heterogeneous computing power resources, including general computing power (CPU) and intelligent computing power (GPU, FPGA, ASIC, etc.). The intelligent computing center covers facilities, hardware, and software, and can provide full-stack capabilities from underlying computing power to top-level application enabling.
[0061] The "intelligent computing center" described in the present invention includes but is not limited to intelligent computing centers.
[0062] The "intelligent computing center" described in the present invention, namely the artificial intelligence computing center, is a type of computing power infrastructure based on artificial intelligence theory, adopting an artificial intelligence computing architecture, and providing computing power services, data services, and algorithm services required for artificial intelligence applications.
[0063] The "computing power center" described in the present invention refers to: a facility mainly composed of infrastructure such as wind, fire, water, and electricity and IT software and hardware devices, with computing power, transportation power, and storage power, including general data centers, intelligent computing centers, supercomputing centers, etc.
[0064] The "supercomputing center" described in the present invention refers to: namely the supercomputing data center, which is a data center based on supercomputers or large-scale computing clusters, capable of providing functions such as large-scale computing, storage, and network services, and is widely used in application scenarios such as aerospace, national defense, oil exploration, climate modeling, and genome sequencing.
[0065] The "computing power resources" described in the present invention refers to: technologies and facilities required for the development of the digital society with information computing, transmission, storage, and application capabilities, including but not limited to computing resources such as CPU and GPU, network resources such as switches and routers, storage resources such as storage arrays and distributed storage, security resources such as firewalls and intrusion detection systems, and support and guarantee resources such as wind, fire, water, and electricity.
[0066] The "large language model" described in the present invention refers to the large language model (LLM), which is a language model with a relatively large number of parameters, aiming to understand and generate human language, trained through a large amount of text data, and can perform a wide range of tasks including text summarization, translation, sentiment analysis, etc.
[0067] The "Multimodal Large Models" described in the present invention refer to models that jointly train multimodal information such as text, images, videos, and audio, including but not limited to multimodal large language models.
[0068] In the field of computer vision object detection, the fast detection algorithm You Only Look Once (YOLO) series of models are widely used in various practical scenarios due to their end-to-end training method and real-time detection performance. Based on the continuous iteration of YOLO, the detection speed and accuracy in various application scenarios have been continuously improved. In addition, deep learning models for image segmentation, such as the Segment Anything Model (SAM), also show generality and robustness in image segmentation and object extraction. The vision-language hybrid strategy they adopt can quickly locate relevant image regions or objects according to human natural language instructions.
[0069] However, in order to obtain more accurate detection results in object detection, it is usually necessary to obtain a large amount of ground truth with accurate bounding box or segmentation mask annotations. When obtaining these data, traditional detection models use manual annotation to label objects one by one, which not only requires annotators to have good professional levels but also takes a lot of time. In addition, the method of annotation through other platforms also has the problem of uneven quality, which seriously affects the detection accuracy of the model when there are many errors in the annotation.
[0070] The present invention uses a general language-vision model (such as LangSAM) to automatically segment the target area according to the natural language query given by the user, and at the same time converts the obtained segmentation boundary into the Bounding Box format commonly used in object detection and directly writes it into the label file required by YOLO, which can greatly reduce the workload of manual annotation.
[0071] See Figure 1 , Figure 1 is a method for implementing image detection based on the computing power of an intelligent computing center provided by the present invention, including:
[0072] Step S1: Obtain a first image;
[0073] Step S2: Identify the first region where the first object indicated by the input first keyword is located in the first image;
[0074] Step S3: Obtain the information of the first region, and convert the information of the first region into the first region information in a target format, where the target format is a format recognizable by the target detection model;
[0075] Step S4: Establish a correspondence between the first category information to which the first object belongs and the first region information;
[0076] Step S5: Use the first category information, the first region information, and the correspondence to train the target detection model, where the target detection model is used to implement image detection based on the computing power of the intelligent computing center.
[0077] Among them, the first image can be obtained through various methods such as local retrieval, network download, or user upload, which is not limited here.
[0078] The first keyword can be a keyword input by the user, and the first keyword can include information about the first object to be detected. For example, it includes the name of the object, the characteristics of the object, or other relevant information. For example, the keywords are "person", "car", "dog", etc.
[0079] The first keyword can be input by the user before obtaining the first image, or after obtaining the first image. The first keyword and the first image can also be input simultaneously, which is not limited here.
[0080] For example, according to the first keyword "capacitor" input by the user, retrieve and obtain the first image containing the capacitor locally, and identify the specific region where the capacitor is located in the first image. Another example is to identify the specific region where the capacitor is located in the first image according to the first keyword "capacitor" input by the user and the previously obtained first image containing the capacitor.
[0081] The above information of the first region can be obtained through a language-vision model. For example, the user inputs the keyword "dog" of the first object to be detected, and through the language-vision model, multi-modal feature fusion is performed between the visual feature extractor and the text feature encoder to determine the corresponding region or mask of the dog in the image to be detected.
[0082] Based on the first region of the first object in the first image, the information of the first region can be obtained, including the coordinate positions of the feature points of the first object in the image, or the bounding box of the first object, the object region mask, and can also include the size information of the first object in the image, etc. Based on the information of the first region, the position of the first object in the first image can be determined.
[0083] Based on the information of the first region, the first region information of the first object can be determined. For example, based on the boundary region of the first object, the minimum bounding rectangle (Bounding Box, BBox) of the boundary region is obtained. The information of the minimum bounding rectangle (i.e., the first region information) is format-converted so that the target detection model can recognize the first region information, and thus the position of the first object in the first image can be determined based on the recognized first region information. For example, when the target detection model is YOLO, the format of the minimum bounding rectangle can be converted into a pattern recognizable by YOLO. In actual detection, the target detection model can also be other detection models, and the target format can be transformed accordingly according to the detection model.
[0084] A corresponding relationship is established between the obtained first region information and the category information of the first object. Based on this corresponding relationship, the category of the first object and its position in the first image can be obtained, thus completing the process of automatic annotation. The target detection model can use the above corresponding relationship and information for model training. Among them, the category information of the first object can be used to classify the recognized object, and this category information can be the same as or different from the above keywords. The category information is, for example, "dog", "capacitor".
[0085] In some embodiments, the above category information can be keywords input by the user. The keywords input by the user are used as the category information. In this way, when outputting the detected object, the keywords are output, which is convenient for the user to quickly determine the object to be detected.
[0086] In actual operation, the above category information can include the major category for determining the category, such as "capacitor". To improve the accuracy of search, the above classification information can also include the sub-categories obtained by further subdividing the major category, so that when inputting keywords, the keywords corresponding to the corresponding category information can be input. For example, when the user labels the corresponding object by inputting the keyword "damaged capacitor", "damaged capacitor" can be stored as the object category. In this way, when performing image detection, object detection can also be performed by inputting "damaged capacitor", which can improve the detection efficiency.
[0087] In the above process, the computing power of the intelligent computing center is used for automatic annotation, reducing manual operations, improving the annotation efficiency, and controlling the detection accuracy through model adjustment, thereby improving the accuracy of image detection.
[0088] The target detection model can be trained based on the above information. Before training, the system automatically generates a configuration file, such as data.yaml, and writes information such as the directory paths of the training set and the validation set, the total number of target categories, and the list of category names. Manual editing is not required in this process, and the operation is simplified by using the computing power of the intelligent computing center.
[0089] During training, the system can load training weights, that is, use the weight information of a pre-defined and modifiable base model (e.g., YOLO). This pre-training is performed by training on a large number of general datasets and has strong general-level initial feature extraction capabilities. In addition, parameter settings are made for the model, including but not limited to the number of training epochs, the Graphics Processing Unit (GPU) device ID used, the batch size, etc. During the training process, the training loss or the mean Average Precision (mAP) metric of the stage verification is monitored and displayed in real time. When the preset conditions are met (e.g., the preset number of iterations or the early stopping strategy conditions, etc.), the "optimal model weights" are saved for inference use.
[0090] After training is completed, metrics such as Average Precision (AP) and average recall rate are output on the validation set, which is convenient for users to evaluate the performance of the model in the target domain scenario. If the performance fails to meet expectations, it can be retried after adjusting the data source, various thresholds, or the training epoch (Epoch) according to the specific situation.
[0091] Currently, general large models (e.g., SAM) can automatically complete the operations of segmenting or detecting objects from the semantic and visual levels. However, most traditional processes cannot integrate these general large models well with known small model object detection frameworks (such as YOLO), resulting in a low degree of automation in data annotation. By training the object detection model in the above manner of the present invention, making full use of the computing power of the intelligent computing center and the cooperation between the general large model and the traditional detection model, the data is automatically annotated, improving the annotation efficiency. And through the above method, the trained object detection model can be obtained quickly. Based on the computing power of the intelligent computing center, using the object detection model for image detection can improve the detection accuracy of the image.
[0092] Optionally, in some embodiments, after step S5, the method further includes:
[0093] Step S6: Input the second image to be detected and the second keyword into the trained object detection model, where the second keyword is used to indicate the first object;
[0094] Step S7: Use the trained object detection model to detect the category of the first object indicated by the second keyword and the second region information of the first object in the second image, and output the category information of the first object and the second region information.
[0095] Among them, the second keyword and the first keyword can be the same or different, and both the first keyword and the second keyword can be used to indicate the first object. When the first keyword and the second keyword are different, for example, the first keyword and the second keyword are "camera" and "photographic camera" respectively, both can be used to indicate the same object.
[0096] Since the target detection model identifies and labels the first object indicated by the first keyword during training, in the case of obtaining the second keyword, feature extraction can be performed on the second image, and the first object corresponding to the second keyword can be obtained in the second image, and the category information of the first object and the second region information of the region where the first object is located are output.
[0097] Among them, the second region information may be the information of the minimum bounding rectangle (Bounding Box, BBox) of the boundary contour of the first object, and may also be the boundary contour information or mask of the first object.
[0098] For example, the user inputs the image to be detected and inputs the keyword "dog". The target detection model identifies the features of the image, regresses the bounding box on the feature map according to the image features and obtains the category score of the category belonging to "dog", then adjusts the size and position of the bounding box, and after removing duplicate detection results, outputs the position of the bounding box of "dog" in the image and the category information to which the object belongs (this category information can be the keyword "dog").
[0099] When multiple first objects are detected, the multiple first objects are output in sequence.
[0100] Based on the computing power of the intelligent computing center, the output of the high-precision segmentation model is converted into BBox annotation, allowing high-quality annotation accuracy to be maintained in complex scenarios. Compared with the traditional method that completely relies on manual work, automatic annotation can save more than 90% of the annotation time in the scenario of a large number of images, and does not require professional annotators, greatly reducing the probability of misannotation.
[0101] The image detection process in this embodiment and the above image training process can be executed on the same device or on different devices.
[0102] Optionally, in some embodiments, step S7 includes:
[0103] Step S71: Use the trained target detection model to determine the category of at least one first object indicated by the second keyword and at least one second region information of the at least one first object in the second image, and the confidence corresponding to each second region information;
[0104] Step S72: Output the category information of the first object with a confidence level greater than the preset threshold and the corresponding second region information.
[0105] When at least one first object indicated by the second keyword is detected, determine the category to which each first object belongs and the corresponding confidence level respectively.
[0106] To reduce noise and detection error rates, an adjustable confidence threshold, i.e., the preset threshold, can be set in advance. Among all results, only the results with a confidence level higher than this threshold are retained. This threshold can be dynamically adjusted according to the difficulty of object detection and detection accuracy.
[0107] For example, when two regions corresponding to dogs are detected in an image for the keyword "dog" and the confidence levels are displayed. The two confidence levels are 0.9 and 0.2 respectively, and the threshold is 0.6, then only the detection result corresponding to the confidence level of 0.9 (greater than 0.6) is retained and output.
[0108] When there are multiple keywords, multiple detection frames may be obtained, and the regions corresponding to these detection frames can be automatically recorded and integrated.
[0109] Optionally, after the step S5, the method further includes:
[0110] Step S8: Generate an identifier of the object detection model, and establish a mapping relationship between the identifier, the storage path of the object detection model, and the keyword list, where the keyword list includes the keywords input into the object detection model, and the second keyword is included in the keyword list;
[0111] The step S6 includes:
[0112] Step S61: When receiving the input second keyword, determine the object detection model corresponding to the second keyword according to the mapping relationship;
[0113] Step S62: Input the second image to be detected and the second keyword into the trained object detection model.
[0114] After the object detection model is trained, an identifier is generated for the object detection model, and a corresponding relationship is established among the identifier, the keyword list, and the storage path of the object detection model. According to this corresponding relationship, the objects corresponding to the keywords in the keyword list for which the object detection model has been labeled and trained can be determined. When these objects need to be detected, the object detection model can be used for detection without re-labeling and training.
[0115] When a detection request input by the user is received, according to the second keyword input by the user (which belongs to the keywords in the keyword list), and based on the above mapping relationship, the corresponding identifier and the storage path of the target detection model are determined, so that the target detection model can be quickly obtained according to the storage path, and the target detection model is used to detect the object corresponding to the second keyword in the second image, which can improve the detection efficiency. For the same or similar object categories, repeated training is avoided, and the waste of the computing power of the intelligent computing center is reduced.
[0116] Optionally, in some embodiments, the step S1 includes at least one of the following:
[0117] Step S11: Determine the first image associated with the third keyword in the pre-stored images according to the input third keyword and the preset association relationship between the image name and the keyword pre-stored;
[0118] Step S12: Retrieve the first image associated with the fourth keyword in the server according to the input fourth keyword;
[0119] Step S13: Generate the first image associated with the fifth keyword according to the input fifth keyword. Traditional object detection data sets generally rely on public databases or enterprise internal business data, and it is often difficult to cover all object categories and forms. For some enterprises, data is often difficult to collect and relatively scarce, resulting in the model being difficult to obtain good generalization ability.
[0120] The first image in the present invention may include multiple images, and the above multiple images may be obtained respectively from any one or more of the following ways. The first way is to obtain images from the local database.
[0121] By means of a pre-organized CSV file or database project, the keyword is associated with the image file name. Among them, these images may also be user-defined uploaded files. The association relationship between these images and them can be stored in the local existing image library. When the user inputs the third keyword, the system retrieves in sequence according to the association relationship and obtains the first image matching the third keyword, and copies the image from the local library to the temporary training directory. This process can be finely controlled according to the upper limit number of different keywords.
[0122] The second way is to retrieve images from the server.
[0123] Communicate through a remote image search service. For example, send a retrieval request for the fourth keyword and the required number of pictures via a POST request. The server will use protocols such as HTTP to retrieve the corresponding first image from the network in real time from an image retrieval engine and return the corresponding first image path information. The system will then copy the first image from the server or network storage to a temporary directory. It can also be an external link address provided by a picture search engine to obtain an image from a remote image database and store it in a specified directory.
[0124] This part of the data usually has high diversity and real-world scenario distributions. The content retrieved each time is different, and real-time and up-to-date results for certain keywords can be obtained, which can improve the generalization performance of the model.
[0125] The third method is text-to-image generation.
[0126] After a certain number of retrievals, the quality of the pictures starts to decline rapidly. To overcome the problem that it is difficult to obtain and even missing some category samples in real-world data, the system integrates a text-to-image generation model. Through natural language description, using a model (such as SDXL), images related to keywords are synthesized within a certain time to enrich the training set. In addition to expanding the data, synthetic data can also avoid potential copyright or privacy disputes. The above-mentioned first image can be obtained based on the fifth keyword.
[0127] In the above-mentioned multiple methods, the data sources can be increased or decreased according to the application scenario. For example, more picture download services can be accessed according to the scenario, or switched to the local picture library mode only when there is no network in the customer scenario. Through the above-mentioned multiple methods, diverse image data can be obtained, enriching the distribution and diversity of the training data, and solving the situation where it was difficult to train in the past due to insufficient materials or inaccurate annotations provided by users. At the same time, users can also manually upload datasets, and the mixing of various data sources improves the quantity and diversity of the training data.
[0128] Optionally, in some embodiments, step S3 includes:
[0129] Step S31: Obtain the pixel mask or region boundary contour of the first region, and determine the minimum bounding rectangle of the region corresponding to the pixel mask or the region boundary contour to obtain the information of the minimum bounding rectangle;
[0130] Step S32: Convert the format of the information of the minimum bounding rectangle to the target format, and perform normalization processing on the information of the minimum bounding rectangle. After the normalization processing, the information of the minimum bounding rectangle is within the target interval, and the first region information in the target format includes the information of the minimum bounding rectangle after the normalization processing.
[0131] When the information of the first region includes the pixel mask of the first region, determine the minimum bounding rectangle of the pixel mask; when the information of the first region includes the regional boundary contour, determine the minimum bounding rectangle of the regional boundary contour (i.e., the information of the first region).
[0132] Convert the information of the first region into the information of the first region in the above manner. The information of the first region includes the information of the minimum bounding rectangle. In order to enable the target detection model to recognize this information, convert the information of the minimum bounding rectangle into a format.
[0133] For example, when using YOLO as the target detection model, convert the segmented pixel mask or coordinate contour into a minimum bounding rectangle (Bounding Box, BBox), and then further convert it into data in the format required by YOLO (center point x, center point y, width, and height) (i.e., the target format), and normalize these data to the interval [0, 1] (i.e., the target interval). Finally, record the target category index and the BBox in the label file, and multiple target BBoxes of an image will be written line by line after normalization, thus completing automatic annotation.
[0134] Taking YOLO as the target detection model as an example above, when using other detection models, it can be adjusted accordingly to the format and the information of the first region that the target detection model can recognize.
[0135] In this way, convert the output of the high-precision segmentation model into BBox annotation, allowing high-quality annotation accuracy to be maintained in complex scenarios.
[0136] In the present invention, through automatic multi-source data collection and text-to-image generation, even in the absence of image data for new target categories, a training set of any scale but with diversity can be constructed in a short time. Combining with a language-vision segmentation model to complete automatic annotation, without the intervention of experienced personnel, saving time and labor costs. And the automatic connection of the training process and the inference process can obtain a new detection model with a certain accuracy in a short time.
[0137] The above detection method of the present invention can be completed based on the following modules of the system.
[0138] 1. Data automatic retrieval module. It includes the following sub-modules:
[0139] Local retrieval sub-module: Match the corresponding picture resources from the local image library according to the user's natural language, and copy or link them to a temporary training folder. It can support the establishment and use of the local picture library or allow the user to upload their own pictures.
[0140] Server Retrieval Sub-module: Communicate with the specified server via protocols such as HTTP or HTTPS according to the keywords matched by the user's natural language.
[0141] Text-to-Image Generation Sub-module: Adopt a text-to-image generation model (such as SDXL) to generate relevant image materials according to the input natural language description for augmenting the training dataset for the same target keyword.
[0142] 2. Automatic Annotation and Data Preprocessing Module
[0143] Language-Visual Model Inference Sub-module: Utilize a general model (such as LangSAM) to automatically locate or segment relevant regions that are associated and conform to the target domain definition in the obtained images through the given keywords.
[0144] Bounding Box Conversion Sub-module: Since the above region results are only image masks, it is necessary to convert the masks into the Bounding Box annotation format with weights commonly used in object detection models (such as YOLO) and write them into the corresponding label files in the order of images.
[0145] Data Configuration Sub-module: Automatically generate the configuration files (such as data.yaml) required for object detection model training, including key training information such as the paths of the training set and validation set and the list of target categories according to the required structure.
[0146] 3. Model Training and Evaluation Module
[0147] Model Initialization Sub-module: Load the pre-trained object detection model (such as YOLOv8 or other models) and manage it uniformly using the model manager.
[0148] Training Scheduling Sub-module: Coordinate the underlying DL framework (such as PyTorch) based on key training parameters such as the given number of training epochs, device list (GPU / CPU), and batch size to complete the training process.
[0149] Training Log and Model Evaluation Sub-module: Monitor the training progress and Loss curve in real time during the training process, and output key evaluation metrics such as the optimal model weights and mAP after the training ends.
[0150] 4. Inference and Model Registration Module
[0151] Inference Sub-module: When the model training is completed or an existing model needs to be used, load the trained weights for inference and generate prediction results (including target category and bounding box information).
[0152] Model version management sub-module: For a newly trained model, a unique identifier is automatically generated and the mapping relationship between it and the corresponding target category is written into the configuration file for subsequent quick retrieval and use of known categories in related tasks.
[0153] The inference and model registration module can support model rollback and differential management.
[0154] To facilitate the understanding of the present invention, the following takes the detection of "component defects" as an example for illustration.
[0155] 1. Determine the target category and keywords
[0156] Through the pre-planner or the developer-specified detection keyword list, such as "capacitor", "capacitor damage", "capacitor deformation", etc.
[0157] 2. Data collection
[0158] Search for images in the local image library to find some images with capacitors and their damage or deformation.
[0159] Batch retrieve circuit board images in industrial scenarios from the image search server, focusing on the locations where capacitors and their parts are damaged or deformed.
[0160] Use the generative model to generate synthetic images according to the description of "damaged capacitor" to further expand the sample diversity.
[0161] 3. Automatic annotation
[0162] Use the multi-modal annotation model to automatically segment and annotate the collected images, filter the bounding boxes BBox of the objects through the confidence threshold, and convert them into the YOLO training format.
[0163] 4. Configuration and training
[0164] The system automatically generates a data.yaml file and writes the number of target categories and category names. Subsequently, load the pre-trained YOLO model for real-time training.
[0165] 5. Inference and verification
[0166] After training is completed, several new circuit board images containing capacitor damage or deformation can be selected for inference, and the capacitors and their defects are detected and annotated by the model.
[0167] 6. Model registration
[0168] The optimal model weights are written into the system configuration file according to the correspondence between the Universally Unique Identifier (UUID) and keywords such as "capacitor, capacitor damage, capacitor deformation", laying a foundation for subsequent automatic calls.
[0169] The present invention combines multiple models such as a language-vision segmentation model, an object detection model, and a text-to-image generation model to achieve automatic annotation and detection of images. The present invention has the following beneficial effects:
[0170] In zero-shot or few-shot scenarios, in the case of lack of user's original pictures or annotations (no annotations or few annotations), it can quickly detect and identify specific targets; without a large number of professional manual annotations, it automatically completes various configurations of annotation generation and model training; integrates multi-source data (obtains image data through various channels); has good versatility and scalability, and can add new target categories or use other detection frameworks in new tasks; improves annotation efficiency.
[0171] The above solution of the present invention can be applied to various fields that require rapid training of customized object detection, such as industrial inspection, medical imaging, product retail, factory security, and driverless driving. It has a high degree of automation, relatively high detection accuracy, and flexible data processing methods, and can be applied to various scenarios. Compared with traditional object detection and annotation methods, while simplifying and automating the data acquisition and annotation process, it uses existing models and computing power to greatly reduce the training time and computing power cost.
[0172] See Figure 2 , Figure 2 which is a device for realizing image detection based on the computing power of an intelligent computing center provided by the present invention. The device includes:
[0173] An acquisition module 201, configured to acquire a first image;
[0174] An identification module 202, configured to identify a first region where a first object indicated by a first keyword input in the first image is located;
[0175] A conversion module 203, configured to acquire information of the first region and convert the information of the first region into first region information in a target format, where the target format is a format recognizable by an object detection model;
[0176] A establishment module 204, configured to establish a correspondence between first category information to which the first object belongs and the first region information;
[0177] A training module 205, configured to train the object detection model by using the first category information, the first region information, and the corresponding relationship, where the object detection model is used to implement image detection based on the computing power of the intelligent computing center.
[0178] Optionally, the apparatus further includes:
[0179] An input module, configured to input a second image to be detected and a second keyword into the trained object detection model, where the second keyword is used to indicate the first object;
[0180] An output module, configured to use the trained object detection model to detect the category of the first object indicated by the second keyword and the second region information of the first object in the second image, and output the category information of the first object and the second region information.
[0181] Optionally, the output module includes:
[0182] A first determination sub-module, configured to use the trained object detection model to determine the category of at least one first object indicated by the second keyword, the at least one second region information of the at least one first object in the second image, and the confidence level corresponding to each second region information;
[0183] An output sub-module, configured to output the category information of the first object with a confidence level greater than a preset threshold and the corresponding second region information.
[0184] Optionally, the apparatus further includes:
[0185] A generation module, configured to generate an identifier of the object detection model, and establish a mapping relationship between the identifier, the storage path of the object detection model, and a keyword list, where the keyword list includes keywords input into the object detection model, and the keyword list includes the second keyword;
[0186] The input module includes:
[0187] A second determination sub-module, configured to determine the object detection model corresponding to the second keyword according to the mapping relationship when receiving the input second keyword;
[0188] An input sub-module, configured to input the second image to be detected and the second keyword into the trained object detection model.
[0189] Optionally, the acquisition module is configured to perform at least one of the following:
[0190] Determine the first image associated with the third keyword in the pre-stored images according to the input third keyword and the preset association relationship between the pre-stored image names and keywords;
[0191] Retrieve the first image associated with the fourth keyword in the server according to the input fourth keyword;
[0192] Generate the first image associated with the fifth keyword according to the input fifth keyword.
[0193] Optionally, the conversion module includes:
[0194] An acquisition sub-module, configured to acquire the pixel mask or the regional boundary contour of the first region, and determine the minimum bounding rectangle of the region corresponding to the pixel mask or the regional boundary contour, so as to obtain the information of the minimum bounding rectangle;
[0195] A conversion sub-module, configured to convert the format of the information of the minimum bounding rectangle into the target format, and perform normalization processing on the information of the minimum bounding rectangle, so that the information of the minimum bounding rectangle after normalization processing is within the target interval, and the first region information in the target format includes the information of the minimum bounding rectangle after normalization processing.
[0196] The device for implementing image detection based on the computing power of the intelligent computing center provided by the embodiments of the present invention can implement each process of the above method for implementing image detection based on the computing power of the intelligent computing center. The technical features correspond one by one and can achieve the same technical effects. To avoid repetition, they are not described here again.
[0197] It should be noted that the device for implementing image detection based on the computing power of the intelligent computing center in the embodiments of the present invention can be a device, or a component, an integrated circuit, or a chip in an electronic device.
[0198] The embodiments of the present invention further provide an electronic device. Refer to Figure 3 , Figure 3 is a schematic structural diagram of an electronic device provided by the embodiments of the present invention. The electronic device includes a memory 1601, a processor 1602, and a program or instruction stored on the memory 1601 and running. When the program or instruction is executed by the processor 1602, it can implement Figure 1 any step in the corresponding method embodiment for implementing image detection based on the computing power of the intelligent computing center and achieve the same beneficial effects, which are not described here again.
[0199] Among them, the processor 1602 may be a Central Processing Unit (CPU), an Application-Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or a Graphics Processing Unit (GPU).
[0200] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiment for realizing image detection based on the computing power of an intelligent computing center, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here. Among them, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk, an optical disc, or the like.
[0201] The embodiment of the present invention also provides a computer program product, including computer instructions, which when executed by a processor, implement each process of the above Figure 1 method embodiment for realizing image detection based on the computing power of an intelligent computing center as shown, and can achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0202] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article, or device including that element.
[0203] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0204] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit and scope protected by the claims of the present invention, and all of them belong to the protection scope of the present invention.
Claims
1. A method for realizing image detection based on the computing power of an intelligent computing center, characterized in that: include: Step S1: Acquire a first image; Step S2: identifying in the first image a first area where a first object indicated by an input first keyword is located; Step S3: Acquire the information of the first region, and convert the information of the first region into first region information in a target format, wherein the target format is a format recognizable by a target detection model; Step S4: establishing a correspondence between the first category information to which the first object belongs and the first region information; Step S5: Use the first category information, the first area information and the corresponding relationship to train the target detection model, and the target detection model is used to realize image detection based on the computing power of the intelligent computing center.
2. The method according to claim 1, characterized in that After step S5, the method further includes: Step S6: inputting a second image to be detected and a second keyword into the trained object detection model, wherein the second keyword is used to indicate the first object; Step S7: using the trained object detection model to detect the category of the first object indicated by the second keyword and the second region information of the first object in the second image, and outputting the category information of the first object and the second region information.
3. The method according to claim 2, characterized in that The step S7 comprises: Step S71: using the trained object detection model to determine the category of at least one first object indicated by the second keyword and at least one second region information of the at least one first object in the second image, as well as the confidence corresponding to each second region information; Step S72: Outputting the category information of the first object and the corresponding second region information whose confidence level is greater than a preset threshold.
4. The method according to claim 2 or 3, characterized in that: After step S5, the method further includes: Step S8: generating an identifier of the target detection model, and establishing a mapping relationship between the identifier, a storage path of the target detection model and a keyword list, wherein the keyword list includes keywords input into the target detection model, and the keyword list includes the second keyword; The step S6 comprises: Step S61: when receiving a second keyword input, determining the object detection model corresponding to the second keyword according to the mapping relationship; Step S62: inputting the second image to be detected and the second keyword into the trained target detection model.
5. The method according to any one of claims 1 to 3, characterized in that The step S1 includes at least one of the following: Step S11: determining the first image associated with the third keyword from pre-stored images according to the input third keyword and the pre-stored preset association relationship between the image name and the keyword; Step S12: according to the input fourth keyword, searching the server for the first image associated with the fourth keyword; Step S13: generating the first image associated with the fifth keyword according to the input fifth keyword.
6. The method according to any one of claims 1 to 3, characterized in that The step S3 comprises: Step S31: acquiring a pixel mask or a region boundary contour of the first region, and determining a minimum bounding rectangle of the region corresponding to the pixel mask or the region boundary contour, to obtain information of the minimum bounding rectangle; Step S32: converting the format of the minimum enclosing rectangle information into the target format, and normalizing the minimum enclosing rectangle information, wherein the normalized minimum enclosing rectangle information is within the target interval, and the first area information of the target format includes the normalized minimum enclosing rectangle information.
7. A device for realizing image detection based on the computing power of an intelligent computing center, characterized in that: include: An acquisition module, used for acquiring a first image; a recognition module, configured to recognize, in the first image, a first area where a first object indicated by an input first keyword is located; a conversion module, configured to obtain information of the first region and convert the information of the first region into first region information in a target format, wherein the target format is a format recognizable by a target detection model; An establishing module, used for establishing a corresponding relationship between the first category information to which the first object belongs and the first region information; A training module is used to train the target detection model using the first category information, the first area information and the corresponding relationship, and the target detection model is used to realize image detection based on the computing power of the intelligent computing center.
8. An electronic device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, the steps of the method for realizing image detection based on the computing power of an intelligent computing center as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for realizing image detection based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.
10. A computer program product, characterized in that It comprises computer instructions, which, when executed by a processor, implement the steps of the method for realizing image detection based on the computing power of an intelligent computing center as described in any one of claims 1 to 6.
Citation Information
Cited By
Method and device for automatically detecting injection-writing consistency of automobile drawings by intelligent computing cloud platform through computing power
CN121392886A
Method and device for automatically detecting consistency of automobile drawing annotation by using computing power of intelligent computing cloud platform
CN121392886B