Machine vision deep learning model training method based on sample characteristic distribution diagram, electronic equipment and storage medium
By establishing a spatial distribution map of sample features and correcting the problem sample features, the training problem caused by the unstable sample set of machine vision deep learning models is solved, and better model inference effect and cost savings are achieved.
Patent Information
- Application Number
- CN202510541500.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The existing machine vision deep learning model has become more and more unstable in the sample set of project iterations, resulting in a lot of time and manpower required for sample cleaning and difficult to locate problems.
By obtaining the sample set of material surface defects and machine vision deep learning models, a spatial distribution map of sample features is established, and high-dimensional features are mapped to low-dimensional space using manifold dimensionality reduction technology, positioning and correcting problem sample features, and finally retraining the model.
Based on the visual characteristics of the sample feature spatial distribution map, it can quickly locate problem sample features, save time and labor costs, and improve the model's inference effect.
Smart Images

Figure CN120071050A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision deep learning, and in particular to a method for training a machine vision deep learning model based on a sample feature distribution map, an electronic device, and a storage medium. Background Art
[0002] In the process of implementing a machine vision deep learning project related to material surface defects, sample marking and sample cleaning work are crucial. However, it also requires a large amount of time and manpower, and is prone to errors. Once an error occurs, it will affect the training and inference effects of the model, and it is difficult to troubleshoot in the later stage. Specifically, in the initial stage of the project, the user provides an original sample set, and the feature categories need to be marked according to the original sample set. As the project iterates, new sample features are continuously added to the sample set. Factors such as changes in marking personnel, marking standards, and detection standards accumulated over a long time result in an increasingly unstable sample set. Once the trained model does not meet expectations, cleaning the samples requires a large amount of time and manpower, and it is difficult to locate the problem. Summary of the Invention
[0003] The technical problem to be solved by the present invention is that the existing machine vision deep learning model fails to meet expectations in training due to the increasingly unstable sample set during project iteration, and provides a method for training a machine vision deep learning model based on a sample feature distribution map, an electronic device, and a storage medium.
[0004] To solve the above technical problem, the technical solution of the present invention is as follows: A method for training a machine vision deep learning model based on a sample feature distribution map, including: Step S1, obtaining a sample set regarding material surface defects and a machine vision deep learning model trained based on the sample set; Step S2, using the machine vision deep learning model to inspect the sample set and establishing a sample feature space distribution map according to the inspected feature information results: The distribution point pattern of the sample features in the sample feature space distribution map is determined according to the feature categories determined by the user, and the distribution point positions of the sample features in the sample feature space distribution map are determined by applying a manifold dimensionality reduction technique to map the sample features from a high-dimensional space to a low-dimensional space; Step S3, based on the visualization characteristics of the sample feature space distribution map, locating the sample features with problems and correcting them to obtain the corrected sample set; and, Step S4, retraining the machine vision deep learning model based on the corrected sample set.
[0005] As a preferred solution of the method for training a machine vision deep learning model, the material surface defects are sample surface defects of film materials, sample surface defects of plate materials, or sample surface defects of sheet materials.
[0006] As a preferred solution of the machine vision deep learning model training method, in step S3, in the sample feature space distribution map, if there are some sample features falling within the defined regions of other feature categories, then determine the incorrect one between the machine vision deep learning model and the user and correct the incorrect one.
[0007] As a preferred solution of the machine vision deep learning model training method, in step S3, in the sample feature space distribution map, if the defined region ranges of two feature categories overlap or are close, then correct the feature categories of all sample features inside the defined regions of the two feature categories to the same feature category.
[0008] As a preferred solution of the machine vision deep learning model training method, in step S3, in the sample feature space distribution map, if there are some sample features far from the defined regions of all feature categories, then correct the feature categories of this part of sample features to a new added feature category.
[0009] As a preferred solution of the machine vision deep learning model training method, in step S3, if the number of sample features of the new added feature category is too small, then appropriately expand its quantity.
[0010] As a preferred solution of the machine vision deep learning model training method, in step S2, the distribution point positions of the sample features missed by the machine vision deep learning model in the sample feature space distribution map are determined according to the defined regions where the sample features of the same feature category not missed by the machine vision deep learning model are located, and the distribution point styles of the missed sample features and the not missed sample features are differentiated. In step S3, in the sample feature space distribution map, if the distribution point styles of some sample features are differentiated, then check the reasons for the missed detection of this part of sample features or adjust the marking standard of the machine vision deep learning model for this part of sample features.
[0011] Another technical solution of the present invention is as follows: An electronic device includes a processor and a memory. At least one instruction is stored in the memory. The instruction is loaded and executed by the processor to implement the operations performed by the machine vision deep learning model training method based on the sample feature distribution map.
[0012] Another technical solution of the present invention is as follows: A storage medium stores at least one instruction. The instruction is loaded and executed by a processor to implement the operations performed by the machine vision deep learning model training method based on the sample feature distribution map.
[0013] Compared with the prior art, the beneficial effects of the present invention are at least as follows: Based on the visualization characteristics of the sample feature space distribution diagram, the current state of the sample set can be intuitively obtained, the sample features of the problem can be quickly located, providing a reliable basis for adjusting the sample set, and a large amount of time and labor costs can be saved. At the same time, the machine vision deep learning model trained based on the corrected sample set has better inference effects.
[0014] In addition to the technical problems solved by the present invention, the technical features constituting the technical solutions, and the beneficial effects brought by these technical features of the technical solutions described above, other technical problems that the present invention can solve, other technical features included in the technical solutions, and the beneficial effects brought by these technical features will be further described in detail in conjunction with the accompanying drawings. Brief Description of the Drawings
[0015] Figure 1 is a flowchart of the method for training a machine vision deep learning model according to an embodiment of the present invention.
[0016] Figure 2 is a schematic diagram of a sample set according to an embodiment of the method for training a machine vision deep learning model of the present invention (the feature category is determined by the user).
[0017] Figure 3 is a schematic diagram of a sample feature space distribution diagram according to an embodiment of the method for training a machine vision deep learning model of the present invention.
[0018] Figure 4 is a schematic diagram of a sample feature space distribution diagram according to an embodiment of the method for training a machine vision deep learning model of the present invention (the corrected ideal sample set).
[0019] Figure 5 is a schematic structural diagram of an electronic device according to an embodiment of the method for generating a defect contour of the present invention. Detailed Description of the Embodiments
[0020] The present invention will be further described in detail below in conjunction with the accompanying drawings through specific embodiments. It should be noted here that the description of these embodiments is for helping to understand the present invention, but does not constitute a limitation to the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0021] See Figure 1 , an embodiment of the present application provides a method for training a machine vision deep learning model based on a sample feature distribution diagram. The method for training the machine vision deep learning model includes: Step S1, obtain a sample set regarding surface defects of materials and a machine vision deep learning model trained based on the sample set. Among them, the surface defects of the materials are surface defect samples of film materials, sheet materials or sheet materials. The sample set includes: sample features, and the feature categories of the sample features are determined by the user (that is, the user makes an initial label for the sample features). Among them, the sample features include: original sample features in the project startup phase and new sample features in the project iteration phase. In other embodiments, the sample features may only include the original sample features in the project startup phase and do not include the new sample features in the project iteration phase.
[0022] Step S2, the machine vision deep learning model inspects the sample set and establishes a sample feature space distribution map according to the inspection result of the feature information: the distribution point style of the sample features in the sample feature space distribution map is determined according to the feature categories determined by the user. The distribution point styles of the sample features with the same feature category are the same, and the distribution point styles of the sample features with different feature categories are different. The distribution point positions of the sample features in the sample feature space distribution map are determined by applying manifold dimensionality reduction techniques (including but not limited to t-SNE, UMAP, etc.) to map the sample features from a high-dimensional space to a low-dimensional space.
[0023] The machine vision deep learning model may miss some of the sample features during inspection. Since the missed sample features do not have inspection results of feature information, the distribution point positions of the missed sample features in the sample feature space distribution map are determined according to the defined regions of the same feature category that are not missed by the machine vision deep learning model, and the distribution point styles of the missed sample features are differentiated from those of the non-missed sample features, aiming to make it easier for the user to distinguish which are the missed sample features and which are the non-missed sample features.
[0024] See Figure 2 and Figure 3 , the feature categories of the sample features in the sample set are respectively determined by the user as four categories: A, B, C, and D. Among them, solid dots with different shades of gray are used to represent the distribution point styles of the sample features of each feature category in the sample feature space distribution map. Specifically, the gray level of category A is the lightest, the gray level of category B is the second lightest, the gray level of category C is the second darkest, and the gray level of category D is the darkest. The missed sample features of category A are placed in the defined region of the non-missed sample features of category A, and the distribution point style of the missed sample features of category A is a dotted dot, which is different from the solid dot of the non-missed sample features of category A.
[0025] In step S3, based on the visualization characteristics of the sample feature space distribution diagram, the user or the computer can relatively quickly locate the sample features of the problem and correct them to obtain the corrected sample set.
[0026] In the sample feature space distribution diagram, if the distribution point styles of some sample features are differentiated (i.e., there are cases where the machine vision deep learning model fails to detect), then check the reasons for the missed detection of these sample features or adjust the determination criteria of the machine vision deep learning model for these sample features.
[0027] In the sample feature space distribution diagram, if some sample features fall within the defined regions of other feature categories (i.e., the feature categories marked by the user for these sample features do not match the feature categories detected by the machine vision deep learning model), then determine the wrongdoer between the machine vision deep learning model and the user and correct the wrongdoer. If the wrongdoer is the user, then correct the feature category of these sample features to the feature category determined by the machine vision deep learning model. If the wrongdoer is the machine vision deep learning model, then adjust the determination criteria of the machine vision deep learning model for these sample features.
[0028] In the sample feature space distribution diagram, if the ranges of the defined regions of two feature categories overlap or are close, then correct the feature categories of all sample features within the defined regions of the two feature categories to the same feature category.
[0029] In the sample feature space distribution diagram, if some sample features are far from the defined regions of all feature categories, then correct the feature categories of these sample features to a new feature category. Further, if the number of these sample features is small, then expand the sample features of this new feature category to ensure the balance of the sample features of each feature category.
[0030] See Figure 3 , the sample features shown in area I are the missed detections (the user label is category A). The distribution point style of the missed sample features is a dotted circle, and the distribution point style of the non-missed sample features is a solid circle. At this time, check the reasons for the missed detection of the sample features or adjust the marking criteria of the machine vision deep learning model for these sample features.
[0031] See Figure 3 , the sample features of category B shown in area II fall within the defined region of category D, then correct the feature category of the sample features to category D.
[0032] See Figure 3, if the defined areas of Class B and Class C shown in Region III overlap, then unify the feature categories of the sample features within the defined areas of Class B and Class C.
[0033] See Figure 3 , if the distribution points of some sample features shown in Region IV and Region V are far from the areas of Class A, Class B, Class C, and Class D, then modify the feature categories of these sample features to new feature categories.
[0034] See Figure 3 , if the number of sample features shown in Region IV is too small, then appropriately increase the number of samples in its feature category.
[0035] See Figure 4 , the figure shows the corrected sample set, and re - determine Class A, Class B, Class C, Class D, and Class E.
[0036] Step S4, based on the corrected sample set, retrain the machine vision deep learning model. After verification, the inference effect of the retrained machine vision deep learning model has been significantly improved compared with before.
[0037] See Figure 5 , this embodiment of the present application also provides an electronic device. The electronic device includes: a processor and a memory. The processor is connected to the memory. The processor is used to execute the computer program stored in the memory so that the electronic device executes the machine vision deep learning model training method based on the sample feature distribution map.
[0038] Specifically, the processor can be a general - purpose processor, including a central processing unit (Central Processing Unit, abbreviated as CPU), a network processor (Network Processor, abbreviated as NP), etc.; it can also be a digital signal processor (Digital Signal Processor, abbreviated as DSP), an application - specific integrated circuit (Application Specific Integrated Circuit, abbreviated as ASIC), a field - programmable gate array (Field Programmable Gate Array, abbreviated as FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0039] Specifically, the memory includes: ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc and other various media that can store program codes.
[0040] Optionally, the electronic device further includes: a display. The display is communicatively connected to the memory and the processor respectively. The display is configured to display a relevant graphical user interface (GUI) of the machine vision deep learning model training method based on the sample feature distribution map.
[0041] An embodiment of the present application further provides a storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the machine vision deep learning model training method based on the sample feature distribution map is implemented.
[0042] The above only expresses the embodiments of the present invention, and the description thereof is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A machine vision deep learning model training method based on sample feature distribution graph, characterized in that: include: Step S1, obtaining a sample set of material surface defects and a machine vision deep learning model trained based on the sample set; Step S2, the machine vision deep learning model verifies the sample set and establishes a sample feature space distribution map based on the feature information results obtained by the verification: the distribution point pattern of the sample features in the sample feature space distribution map is determined based on the feature category determined by the user, and the distribution point position of the sample features in the sample feature space distribution map is determined by mapping the sample features from a high-dimensional space to a low-dimensional space using a manifold dimensionality reduction technique; Step S3, based on the visualization characteristics of the sample feature space distribution diagram, locate the sample features of the problem and correct them to obtain the corrected sample set; as well as, Step S4: retrain the machine vision deep learning model based on the corrected sample set.
2. The machine vision deep learning model training method according to claim 1, characterized in that: The material surface defect is a film material surface defect sample, a plate material surface defect sample or a sheet material surface defect sample.
3. The machine vision deep learning model training method according to claim 1, characterized in that: In step S3, if some sample features in the sample feature space distribution map fall within the defined area of other feature categories, the error between the machine vision deep learning model and the user is determined and corrected.
4. The machine vision deep learning model training method according to claim 1, characterized in that: In step S3, in the sample feature spatial distribution map, if there are two feature categories whose defined areas overlap or are similar, the feature categories of all sample features within the defined areas of the two feature categories are corrected to the same feature category.
5. The machine vision deep learning model training method according to claim 1, characterized in that: In step S3, if there are some sample features in the sample feature spatial distribution map that are far away from the defined area of all feature categories, the feature categories of these sample features are corrected to newly added feature categories.
6. The machine vision deep learning model training method according to claim 5, characterized in that: In step S3, if the number of sample features of the newly added feature category is too small, the number is appropriately expanded.
7. The machine vision deep learning model training method according to any one of claims 1 to 6, characterized in that: In step S2, the distribution point positions of the sample features missed by the machine vision deep learning model in the sample feature space distribution map are determined based on the defined area where the sample features of the same feature category that are not missed by the machine vision deep learning model are located, and the distribution point patterns of the missed sample features and the sample features that are not missed are differentiated.
8. The machine vision deep learning model training method according to claim 7, characterized in that: In step S3, if the distribution point patterns of some sample features in the sample feature spatial distribution map are differentiated, the reasons for missed detection of these sample features are checked or the marking standards of the machine vision deep learning model for these sample features are adjusted.
9. An electronic device, comprising a processor and a memory, wherein the memory stores at least one instruction, characterized in that: The instructions are loaded and executed by the processor to implement the operations performed by the machine vision deep learning model training method based on the sample feature distribution map described in any one of claims 1 to 6.
10. A storage medium, wherein at least one instruction is stored in the storage medium, characterized in that: The instructions are loaded and executed by the processor to implement the operations performed by the machine vision deep learning model training method based on the sample feature distribution map as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Sample category label correction method and apparatus, and electronic device
CN110457155A
Sample set optimization method and device, equipment, medium and product
CN115600109A
Image deviation data classification method based on data self-selection and mark self-correction algorithm
CN115761288A
Label shift detection and adjustment in predictive modeling
US20210406598A1
Layered cybersecurity using spurious data samples
US20240283822A1