Mould texture identification method and system based on multi-scale feature fusion
Through the mold texture recognition method of multi-scale feature fusion, combined with the P2 feature layer and detection layer, residual module and CBAM attention mechanism, the problem of difficulty in detail recognition in mold texture detection is solved, and efficient and accurate texture detection is achieved.
Patent Information
- Application Number
- CN202510564623.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-15
AI Technical Summary
Existing mold texture detection technology is difficult to efficiently identify the details of complex textures, resulting in misjudgment and low detection efficiency.
A mold texture recognition method with multi-scale feature fusion is adopted. By combining the P2 feature layer and detection layer of the baseline network, the residual module is constructed and feature map convolution is performed, combined with the CBAM attention mechanism, and a mixed data set is constructed for training and testing.
It realizes fine analysis of mold texture, improves the recognition ability of texture details, and improves the accuracy and efficiency of detection.
Smart Images

Figure CN120495691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of feature recognition, and in particular to a mold texture recognition method and system based on multi-scale feature fusion. Background Art
[0002] In the industrial production of molds, mold defect detection is a critical task in the mold manufacturing field. Furthermore, molds are important processing tools in industrial production, and their quality directly affects product quality. With the rapid development of my country's manufacturing industry, the demand for mold quality is becoming increasingly stringent. Traditional mold inspection methods rely primarily on manual labor, resulting in low efficiency, poor accuracy, and high labor intensity. However, mold visual inspection technology, leveraging computer vision, enables automated, high-precision, and efficient inspection, becoming a key tool for improving production efficiency and product quality in the mold industry.
[0003] Molds for consumer electronics like mobile phones, computers, smart wearables, and home appliances, as well as automotive interior and exterior trim, are becoming increasingly sophisticated, and correspondingly, their textures are becoming increasingly complex. Current visual inspection technology struggles to detect these complex textures, so adapting visual inspection to meet mold texture requirements has become a pressing challenge. Summary of the Invention
[0004] In view of the above-mentioned defects, the purpose of the present invention is to propose a mold texture recognition method and system based on multi-scale feature fusion to solve the problem that mold texture details are difficult to be recognized.
[0005] To achieve this purpose, the present invention adopts the following technical solution: a mold texture recognition method based on multi-scale feature fusion, comprising the following steps:
[0006] Step S1: Combine the P2 feature layer of the baseline network architecture of the recognition model with the detection layer;
[0007] Step S2: construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map;
[0008] Step S3: concatenate the output feature maps along the channel and provide them to the CBAM attention mechanism;
[0009] Step S4: Collecting actual scene data of mold texture and obtaining an open source data set to create a mixed data set, and labeling the mixed data set as the first data set;
[0010] The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
[0011] Preferably, the P2 feature layer is combined with the detection layer in step S1 in the following manner:
[0012] The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
[0013] Preferably, the first residual module performs convolution processing on the feature map as follows:
[0014] Step S21: The first residual module applies a 1*1 convolution operation to the channel dimension of the input feature map to split the feature map into S feature map subsets;
[0015] Step S22: except for the first feature map subset, each feature map subset is subjected to a 3*3 convolution operation;
[0016] Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.
[0017] Preferably, the sizes of the output feature maps in step S3 are set to 4.0, 1.0 and 0.4 respectively, and convolution operations of 80*80, 40*40 and 20*20 are performed respectively.
[0018] Preferably, the image size of the first data set in the recognition model is set to 200*200, the initial learning rate is 0.01, the momentum is 0.937, the weight decay factor is 0.005, and the batch size is 8.
[0019] A mold texture recognition system based on multi-scale feature fusion, using the mold texture recognition method based on multi-scale feature fusion, comprising a combining module, a processing module, a splicing module and a training module;
[0020] The combining module is used to combine the P2 feature layer and the detection layer of the baseline network architecture of the recognition model;
[0021] The processing module is used to construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map;
[0022] The splicing module is used to splice the output feature map along the channel and provide it to the CBAM attention mechanism;
[0023] The training module is used to collect actual scene data of mold texture and obtain open source data sets to create a mixed data set, and label the mixed data set as the first data set;
[0024] The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
[0025] Preferably, the P2 feature layer and the detection layer in the combination module are combined in the following manner:
[0026] The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
[0027] Preferably, the processing module includes a splitting submodule and a convolution submodule;
[0028] The split submodule is used in the first residual module to apply a 1*1 convolution operation to the channel dimension of the input feature map, splitting the feature map into S feature map subsets;
[0029] The convolution submodule is used to perform 3*3 convolution operations on each feature map subset except the first feature map subset;
[0030] Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.
[0031] One of the above technical solutions has the following advantages or beneficial effects: the recognition model of the present invention realizes the refined analysis of mold texture through the four-fold design of "feature fusion-residual enhancement-attention calibration-data drive", so that small texture details can be recognized, further improving the ability of mold texture. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] Figure 1 is a flow chart of an embodiment of the method of the present invention.
[0033] Figure 2 It is a structural diagram of an embodiment of the system of the present invention.
[0034] Figure 3 It is a schematic diagram of the fusion of the P2 feature layer and the detection layer in one embodiment of the present invention. DETAILED DESCRIPTION
[0035] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention.
[0036] In the description of the embodiments of the present invention, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of the technical features indicated. Therefore, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the described features. In the description of the embodiments of the present invention, "plurality" means two or more, unless otherwise specifically specified.
[0037] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, "plurality" means two or more. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0038] like Figures 1 to 3 As shown, a mold texture recognition method based on multi-scale feature fusion includes the following steps:
[0039] Step S1: Combine the P2 feature layer of the baseline network architecture of the recognition model with the detection layer;
[0040] Compared with other detection targets, the texture details of the mold are smaller in size, and the corresponding texture details are also relatively small. If the first data set is directly input into the existing recognition model, the recognition model will easily ignore the texture details, resulting in some damaged parts being unable to be recognized, thus causing misjudgment. For this reason, combining the P2 feature layer of the baseline network architecture with the detection layer can reduce information loss, improve the multi-scale and small target recognition capabilities, improve the accuracy of target detection of various sizes, and further enhance the feature representation capabilities.
[0041] Step S2: construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map;
[0042] Step S3: concatenate the output feature maps along the channel and provide them to the CBAM attention mechanism;
[0043] Step S4: Collecting actual scene data of mold texture and obtaining an open source data set to create a mixed data set, and labeling the mixed data set as the first data set;
[0044] The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
[0045] By collecting texture data from actual production line molds (including real noise and lighting variations) and combining it with open-source datasets (providing diverse texture templates), we construct a hybrid dataset. This data augmentation strategy preserves domain-specific features (such as the reflective properties of specific mold materials) while also introducing cross-domain texture samples, improving the model's robustness and generalization.
[0046] The recognition model of the present invention achieves a refined analysis of mold texture through the four-fold design of "feature fusion-residual enhancement-attention calibration-data-driven", allowing small texture details to be recognized, further improving the ability of mold texture.
[0047] Preferably, the P2 feature layer is combined with the detection layer in step S1 in the following manner:
[0048] The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
[0049] Fusion diagram Figure 3 As shown, in one embodiment, the P2 feature layer can be first spliced with different size layers of the detection layer. At this time, the output feature map can retain the multi-scale information and image details in the first data set, so as to better contribute to subsequent detection. Then the feature map is input into the P2 feature layer again and the detection layer is spliced in sequence. At this time, through multi-layer interactive refinement of features, it is suitable for aligning multiple different texture detail features in the mold, realizing the refinement of the texture features of the mold, and being able to better distinguish different texture details in subsequent detections.
[0050] Preferably, the first residual module performs convolution processing on the feature map as follows:
[0051] Step S21: The first residual module applies a 1*1 convolution operation to the channel dimension of the input feature map to split the feature map into S feature map subsets;
[0052] Step S22: except for the first feature map subset, each feature map subset is subjected to a 3*3 convolution operation;
[0053] Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.
[0054] The first layer of the first residual module uses 1×1 convolution to split the input features into S subsets along the channel dimension. This grouping strategy achieves spatial decoupling of features. Each subset can independently learn a specific type of texture pattern. For example, subset 1 may focus on capturing the low-frequency periodic structure of the mold surface (such as the grid base of the injection mold), subset 2 extracts mid-frequency details (such as the transition area of the mold edge), and subset 3-S further refines the high-frequency components (such as surface scratches, micropores and other random defects). This divide-and-conquer strategy not only reduces computational complexity (compared to full-channel convolution, the parameters are reduced by S times), but also encourages the network to explore complementary texture representations through feature isolation, avoiding information mixing in a single path.
[0055] Except for the first subset that retains the original information, the remaining subsets are all subjected to local feature enhancement through 3×3 convolution. On the one hand, the 3×3 convolution kernel has the best local receptive field and can accurately capture the typical scale characteristics of the mold texture (such as the microstructure spacing is usually in the range of 5-20 pixels); on the other hand, by skipping the convolution operation of the first subset, the high-frequency details in the original input are retained, providing a stable benchmark for the feature enhancement of subsequent subsets. Finally, the output of subset i is added to the output of subset i-1 to form residual learning, which can effectively alleviate the gradient disappearance problem caused by increased depth. At the same time, this also allows deep subsets to not need to repeatedly learn the texture details captured by the shallow layer, but instead focus on modeling higher-order texture associations (such as the relationship between the defect distribution pattern and the mold structure). Ultimately, the detection accuracy of the recognition model is improved.
[0056] Preferably, the sizes of the output feature maps in step S3 are set to 4.0, 1.0 and 0.4 respectively, and convolution operations of 80*80, 40*40 and 20*20 are performed respectively.
[0057] This highlights the importance of high-resolution feature maps for improving the accuracy of small object detection. By using classification loss to evaluate the accuracy of classification predictions, the detected objects are accurately classified into different groups. The position change between the predicted box and the ground-truth box is quantified using coordinate loss.
[0058] Preferably, the image size of the first data set in the recognition model is set to 200*200, the initial learning rate is 0.01, the momentum is 0.937, the weight decay factor is 0.005, and the batch size is 8.
[0059] Since many defects in the mold defect dataset are small in size, and their shapes and patterns are complex, diverse and irregular, and different defect categories are also highly similar, this small image size and response parameter setting has higher resolution and sensitivity to correctly detect and classify these defects.
[0060] A mold texture recognition system based on multi-scale feature fusion, using the mold texture recognition method based on multi-scale feature fusion, comprising a combining module, a processing module, a splicing module and a training module;
[0061] The combining module is used to combine the P2 feature layer and the detection layer of the baseline network architecture of the recognition model;
[0062] The processing module is used to construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map;
[0063] The splicing module is used to splice the output feature map along the channel and provide it to the CBAM attention mechanism;
[0064] The training module is used to collect actual scene data of mold texture and obtain open source data sets to create a mixed data set, and label the mixed data set as the first data set;
[0065] The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
[0066] Preferably, the P2 feature layer and the detection layer in the combination module are combined in the following manner:
[0067] The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
[0068] Preferably, the processing module includes a splitting submodule and a convolution submodule;
[0069] The split submodule is used in the first residual module to apply a 1*1 convolution operation to the channel dimension of the input feature map, splitting the feature map into S feature map subsets;
[0070] The convolution submodule is used to perform 3*3 convolution operations on each feature map subset except the first feature map subset;
[0071] Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.
[0072] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative uses of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0073] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A mold texture recognition method based on multi-scale feature fusion, characterized in that: The steps include: Step S1: Combine the P2 feature layer of the baseline network architecture of the recognition model with the detection layer; Step S2: construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map; Step S3: concatenate the output feature maps along the channel and provide them to the CBAM attention mechanism; Step S4: Collecting actual scene data of mold texture and obtaining an open source data set to create a mixed data set, and labeling the mixed data set as the first data set; The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
2. The mold texture recognition method based on multi-scale feature fusion according to claim 1 is characterized in that: The P2 feature layer and the detection layer are combined in step S1 in the following manner: The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
3. The mold texture recognition method based on multi-scale feature fusion according to claim 1, characterized in that: The first residual module performs convolution processing on the feature map as follows: Step S21: The first residual module applies a 1*1 convolution operation to the channel dimension of the input feature map to split the feature map into S feature map subsets; Step S22: except for the first feature map subset, each feature map subset is subjected to a 3*3 convolution operation; Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.
4. The mold texture recognition method based on multi-scale feature fusion according to claim 1, characterized in that: In step S3, the sizes of the output feature maps are set to 4.0, 1.0, and 0.4, respectively, and convolution operations of 80*80, 40*40, and 20*20 are performed correspondingly.
5. The mold texture recognition method based on multi-scale feature fusion according to claim 1, characterized in that: The image size of the first data set in the recognition model is set to 200*200, the initial learning rate is 0.01, the momentum is 0.937, the weight decay factor is 0.005, and the batch size is 8.
6. A mold texture recognition system based on multi-scale feature fusion, using the mold texture recognition method based on multi-scale feature fusion according to any one of claims 1 to 6, characterized in that: It includes combining module, processing module, splicing module and training module; The combining module is used to combine the P2 feature layer and the detection layer of the baseline network architecture of the recognition model; The processing module is used to construct a first residual module and add the first residual module to the back of the detection layer, wherein the first residual module is used for convolution processing of the input feature map; The splicing module is used to splice the output feature map along the channel and provide it to the CBAM attention mechanism; The training module is used to collect actual scene data of mold texture and obtain open source data sets to create a mixed data set, and label the mixed data set as the first data set; The first data set is divided into a training set, a validation set, and a test set according to a certain ratio, and the recognition model is trained, validated, and tested using the first data set.
7. The mold texture recognition system based on multi-scale feature fusion according to claim 6, characterized in that: The P2 feature layer and the detection layer in the combination module are combined in the following manner: The P2 feature layer sequentially fuses the different size layers of the detection layer, and at the same time, the P2 feature layer is fused with the different size layers of the detection layer respectively.
8. The mold texture recognition system based on multi-scale feature fusion according to claim 6, characterized in that: The processing module includes a splitting submodule and a convolution submodule; The split submodule is used in the first residual module to apply a 1*1 convolution operation to the channel dimension of the input feature map, splitting the feature map into S feature map subsets; The convolution submodule is used to perform 3*3 convolution operations on each feature map subset except the first feature map subset; Except for the first feature map subset and the second feature map subset, each feature map subset is also added to the convolution output of the previous feature map subset.