A method, device, electronic device and storage medium for identifying chip silk-screen information

The ResNet and Transformer models combine the cross attention mechanism and the CLIP text encoder to optimize the chip silk screen information recognition, which solves the problem of low-resolution chip silk screen information recognition efficiency and low accuracy, and achieves efficient and accurate recognition effects.

CN118941816BActive Publication Date: 2025-07-11CENT SOUTH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411040076.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-07-11
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

The prior art recognizes low-resolution chip silkscreen information, and affects the recognition efficiency and accuracy.

Method used

The ResNet and Transformer model are combined with the cross attention mechanism, and the silk screen information recognition model is optimized through feature extraction and masking technology, combined with the CLIP text encoder, and the character sequence is expanded by similar words to optimize the model training process.

Benefits of technology

It improves the recognition efficiency and accuracy of chip silk screen printing information, especially in low-resolution scenarios, reduces the cost and time of manual labeling, and enhances the generalization ability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118941816B_ABST
    Figure CN118941816B_ABST
Patent Text Reader

Abstract

The present application discloses a method, apparatus, electronic device and storage medium for chip silk screen information recognition. The method includes: based on the cross-attention mechanism, fusing visual features and language features to obtain fused features, obtaining a sequence of the fused features, randomly selecting a position in the sequence, and using the visual features before the position as the visual features to be masked; classifying the masked visual features to generate a character sequence, mapping the character sequence into a text vector through the text encoder of the pre-trained model, and obtaining similar words generated by the pre-trained model based on the text vector; training the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information; obtaining the silk screen information recognition result output by the trained silk screen information recognition model. The present application is beneficial to improving the recognition accuracy of the silk screen information recognition model for chip silk screen information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and in particular, to a method, apparatus, electronic device, and storage medium for identifying chip silk screen information. Background Art

[0002] In the field of intelligent chip manufacturing, different models and batches of chips can be quickly identified and classified through chip silk screen information, improving production efficiency and automation levels, and contributing to efficient inventory management and production scheduling.

[0003] However, the existing technologies mainly use text detection and segmentation technologies to detect the position information of each character in the chip image, segment them in sequence according to the position information of each character, and finally query and match the characters through manually defined templates to predict and generate a single predicted character. When facing low-resolution chip images, since the character shapes in the chip images are usually relatively blurred and there is a certain degree of image occlusion between characters, it will not only interfere with the accuracy of text detection, but also affect the classification and prediction generation of characters in recognition. Therefore, the existing process for identifying chip silk screen information is cumbersome and not conducive to improving the recognition efficiency and accuracy of chip silk screen information. Summary of the Invention

[0004] Embodiments of this application provide a method, apparatus, electronic device, and storage medium for identifying chip silk screen information to solve the technical problem that the existing process for identifying chip silk screen information is cumbersome and not conducive to improving the recognition efficiency and accuracy of chip silk screen information.

[0005] In a first aspect, embodiments of this application provide a method for identifying chip silk screen information, which is applied to an electronic device. The electronic device stores a silk screen information recognition model, and the silk screen information recognition model includes a ResNet model and a Transformer model. The method for identifying chip silk screen information includes:

[0006] Obtain a preset chip image and preset text description information corresponding to the preset chip image;

[0007] Use the ResNet model and the Transformer model to extract features from the preset chip image to obtain visual features corresponding to the preset chip image, and use the Transformer model to extract features from the preset text description information to obtain language features corresponding to the preset chip image;

[0008] Based on the cross-attention mechanism, fuse the visual features and the language features to obtain fused features, obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked;

[0009] Classify the masked visual features to generate a character sequence, map the character sequence into a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector;

[0010] Optimize the silk screen information recognition model through the CLIP text encoder and the similar words;

[0011] Use the optimized silk screen information recognition model to obtain predicted silk screen information, and train the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information;

[0012] When the cross-entropy loss value is less than the preset loss value, obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0013] In a possible implementation manner of the first aspect, based on the cross-attention mechanism, fuse the visual feature and the language feature to obtain a fused feature, obtain a sequence of the fused feature, randomly select a position in the sequence, and use the visual feature before the position as the visual feature to be masked, including:

[0014] Based on the cross-attention mechanism, fuse the visual feature and the language feature to obtain the fused feature, obtain a sequence of the fused feature, and randomly select a position in the sequence;

[0015] Obtain an attention score map, and based on the attention score map, use the visual feature before the position as the visual feature to be masked.

[0016] In a possible implementation manner of the first aspect, the classifying the masked visual features to generate a character sequence, mapping the character sequence into a text vector through the text encoder of the pre-trained model, and obtaining similar words generated by the pre-trained model based on the text vector includes:

[0017] Replace the feature value of the visual feature to be masked with a selected token to obtain the masked visual feature;

[0018] Send the masked visual feature into a linear layer, and classify the masked visual feature through the classification function to generate a character sequence corresponding to the masked visual feature;

[0019] Map the character sequence into a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector.

[0020] In a possible implementation of the first aspect, optimizing the silk screen information recognition model through the CLIP text encoder and the similar words includes:

[0021] Encode the similar words through the CLIP text encoder to obtain the text features corresponding to the similar words, use the cosine similarity formula to calculate the similarity between the visual features and the text features, and construct a similarity matrix through multiple similarities;

[0022] In the similarity matrix, regard the visual features and the text features with similarity higher than the preset value as positive sample pairs, and regard the visual features and the text features with similarity not higher than the preset value as negative sample pairs;

[0023] Use the contrastive loss function to optimize the silk screen information recognition model with the goal of minimizing the distance between the positive sample pairs and maximizing the distance between the negative sample pairs.

[0024] In a possible implementation of the first aspect, when the cross-entropy loss value is less than the preset loss value, obtaining the silk screen information recognition result output by the trained silk screen information recognition model includes:

[0025] When the cross-entropy loss value is less than the preset loss value, stop training the silk screen information recognition model and save the trained silk screen information recognition model;

[0026] Obtain the recognition accuracy of the trained silk screen information recognition model on the validation set;

[0027] When the recognition accuracy is greater than the preset accuracy, obtain the current chip image and the current text description information corresponding to the current chip image, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0028] In a possible implementation of the first aspect, after obtaining the silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than the preset loss value, the chip silk screen information recognition method includes:

[0029] Create a display window and display the current text description information in the display window.

[0030] In a possible implementation of the first aspect, the pre-trained model includes one or a combination of the Word2Vec model and the BRET model.

[0031] Second aspect, an embodiment of the present application provides a chip silk screen information recognition device, which is applied to an electronic device. The electronic device stores a silk screen information recognition model, and the silk screen information recognition model includes a ResNet model and a Transformer model, including:

[0032] A first acquisition module, configured to acquire a preset chip image and preset text description information corresponding to the preset chip image;

[0033] An extraction module, configured to use the ResNet model and the Transformer model to extract features from the preset chip image to obtain visual features corresponding to the preset chip image, and use the Transformer model to extract features from the preset text description information to obtain language features corresponding to the preset chip image;

[0034] A selection module, configured to fuse the visual features and the language features based on a cross-attention mechanism to obtain fused features, obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked;

[0035] A second acquisition module, configured to classify the masked visual features to generate a character sequence, map the character sequence to a text vector through a text encoder of a pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector;

[0036] An optimization module, configured to optimize the silk screen information recognition model through a CLIP text encoder and the similar words;

[0037] A third acquisition module, configured to use the optimized silk screen information recognition model to obtain predicted silk screen information, and train the silk screen information recognition model according to a cross-entropy loss value between the predicted silk screen information and the true silk screen information;

[0038] A recognition module, configured to obtain a silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than a preset loss value.

[0039] Third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the chip silk screen information recognition method in the first aspect is implemented.

[0040] Fourthly, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the chip silk screen information recognition method in the first aspect above.

[0041] Fifthly, an embodiment of the present application provides a computer program product, which when running on an electronic device, enables the electronic device to execute the chip silk screen information recognition method in the first aspect above.

[0042] The beneficial effects of the embodiments of the present application are in two aspects. On the one hand, when the cross-entropy loss value is less than a preset loss value, the silk screen information recognition result output by the trained silk screen information recognition model is obtained. Since there is no need to manually define a template, the recognition time of chip silk screen information is reduced, which is beneficial to improving the recognition efficiency of chip silk screen information. On the other hand, through the feature masking technology, the visual features are masked, and the ability of the silk screen information recognition model to recognize silk screen information in limited visual features is trained. And by expanding similar words, multiple possible candidates are provided for each predicted character sequence, improving the prediction ability of the silk screen information recognition model in a low-resolution scenario, and further improving the recognition accuracy of the silk screen information recognition model for chip silk screen information. Description of the Drawings

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0044] Figure 1 It is an application scenario diagram of the chip silk screen information recognition method provided by the embodiment of the present application;

[0045] Figure 2 It is a schematic flowchart of the chip silk screen information recognition method provided by the embodiment of the present application;

[0046] Figure 3 It is a flowchart of obtaining the current text description information provided by the embodiment of the present application;

[0047] Figure 4 It is a schematic block diagram of the chip silk screen information recognition device provided by the embodiment of the present application;

[0048] Figure 5 It is a schematic structural diagram of the electronic device provided by the embodiment of the present application. Detailed Embodiments

[0049] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the protection scope of the present application.

[0050] In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present application.

[0051] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all content and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined or partially merged, so the actual execution order may be changed according to the actual situation.

[0052] The chip silk screen information recognition method provided in the embodiments of the present application can be applied to electronic devices such as mobile phones, tablet computers, wearable devices, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPCs), netbooks, personal digital assistants (PDAs), etc. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0053] Please refer to Figure 1 , Figure 1 , which is an application scenario diagram of the chip silk screen information recognition method provided in the embodiments of the present application, and is described in detail as follows:

[0054] The electronic device connects to the database through a preset network, and obtains a preset chip image and preset text description information corresponding to the preset chip image from the database.

[0055] Among them, the preset network includes one or a combination of an Ethernet network, a WIFI network, a 4G network, and a 5G network.

[0056] Among them, the WIFI network is a wireless fidelity network.

[0057] Among them, the 4G network is a fourth-generation mobile communication technology network.

[0058] Among them, the 5G network is a fifth-generation mobile communication technology network.

[0059] In the embodiments of the present application, the electronic device can be connected to a database, and obtain the labeled data of the target chip from the database, which is beneficial to improving the acquisition efficiency of the labeled data.

[0060] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of the chip silk screen information recognition method provided by the embodiments of the present application. This method can be applied to an electronic device, and the electronic device stores a silk screen information recognition model, and the silk screen information recognition model includes a ResNet model and a Transformer model.

[0061] Among them, the ResNet model is a model based on ResNet.

[0062] Among them, ResNet (Residual Network) is a deep neural network architecture.

[0063] Among them, the Transformer model is a neural network model based on the self-attention mechanism.

[0064] Among them, the silk screen information recognition model is a recognition model for silk screen information.

[0065] As Figure 2 shown, the chip silk screen information recognition method provided by the embodiments of the present application includes the following steps, which are described in detail as follows:

[0066] S201, obtain a preset chip image and preset text description information corresponding to the preset chip image;

[0067] Exemplarily, obtaining a preset chip image and preset text description information corresponding to the preset chip image includes:

[0068] Obtain the sample data of the data set, and in the sample data, obtain the preset chip image and the preset text description information corresponding to the preset chip image.

[0069] Among them, before obtaining the sample data of the data set and obtaining the preset chip image and the preset text description information corresponding to the preset chip image in the sample data, the chip silk screen information recognition method includes:

[0070] Use a camera to take multiple angles and views of different models of chips under different backgrounds and lighting conditions to ensure that various possible industrial production scenarios are covered, so as to ensure the integrity of the collected images;

[0071] According to various tag definitions in the tag system, label the silk-screen content in the preset chip image with appropriate silk-screen information tags, and at the same time combine the silk-screen information tags and their coordinate positions into a preset text description information;

[0072] Take the preset chip image and the preset text description information corresponding to the preset chip image as a sample data, and form a data set with multiple sample data.

[0073] S202, use the ResNet model and the Transformer model to extract features from the preset chip image to obtain the visual features corresponding to the preset chip image, and use the Transformer model to extract features from the preset text description information to obtain the language features corresponding to the preset chip image;

[0074] Among them, preprocess the preset chip image and the preset text description information, use the ResNet model and the Transformer model to extract features from the preprocessed preset chip image, and use the Transformer model to extract features from the preprocessed preset text description information to obtain the language features corresponding to the preset chip image.

[0075] Among them, preprocessing the preset chip image and the preset text description information includes:

[0076] Perform various transformations on the preset chip image to generate image pairs with different perspectives and variations to improve the generalization ability of the model;

[0077]

[0078] Among them, x i represents the original image, T represents the image enhancement operation, represents the image after image enhancement.

[0079] The image enhancement operations are as follows:

[0080] Rotation: Rotate the image randomly at an angle between 0 and 360 degrees to simulate different perspectives.

[0081] Scaling: Randomly scale the ratio between 0.8 and 1.5 to generate images of different sizes.

[0082] Flipping: Perform random flipping in the horizontal and vertical directions.

[0083] Cropping: Randomly crop different regions of the image to simulate the loss of partial perspectives;

[0084] Among them, for the text description information, first perform sentence segmentation, then perform word segmentation, stop word removal, and part-of-speech filtering on each sentence. For word segmentation and part-of-speech filtering, use a word segmentation tool, and for stop word removal, use the Harbin Institute of Technology stop word list.

[0085] Among them, the ResNet model and the Transformer model are used to extract the visual sequence features of the preset chip image. Specifically, a two-dimensional convolution is performed on the chip image. While reducing the image size, the preset chip image is mapped to a higher-dimensional space to facilitate the reading of deep features.

[0086] After multiple convolution operations, the input result and the result after convolution will be added together to obtain the image features of the chip image, thereby alleviating the problems of gradient disappearance and explosion caused by too many model layers.

[0087] After passing through the ResNet module, preliminary visual features are obtained. Some pixel points in the feature map are randomly occluded with a certain random probability. The random probability is 0.65 - 0.85, and their values are set to 0. By introducing the two random probabilities, the recognition ability of the model with insufficient visual features is trained, thereby improving the generalization ability of the model.

[0088] In order to extract the sequence features of the preset chip image, the size of the occluded visual features is transposed, merged into three-dimensional channels, and kept consistent with the input dimension of the Transformer. The input is sent to the encoder of the Transformer for feature extraction of the image sequence.

[0089] For visual features, in order to ensure the sequentiality of the character sequence in the preset chip image, the position encoding of the traditional Transformer is adopted, and position encoding is generated for each pixel point through sine and cosine functions.

[0090] After position encoding, the encoded visual features are obtained and sent into the encoder for sequence feature extraction, as shown in the formula. In order to accelerate the training speed of the model, multi-head attention is used for parallel calculation of attention scores, and the final attention scores are obtained through splicing.

[0091] In order to ensure the stability of training and accelerate the training speed, a normalization operation is performed on the encoded visual features, so that the data is distributed within a specific range. Then the normalized data is sent into the feed-forward neural network, enabling the feed-forward neural network to capture more complex dependencies in the data and improving the generalization prediction ability of the silk screen information recognition model.

[0092] After passing through multiple encoders of the Transformer, the silk screen information recognition model obtains deeper visual sequence features.

[0093] S203, based on the cross-attention mechanism, fuse the visual features and the language features to obtain fused features, acquire the sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked;

[0094] Among them, the step of based on the cross-attention mechanism, fusing the visual features and the language features to obtain fused features, acquiring the sequence of the fused features, randomly selecting a position in the sequence, and using the visual features before the position as the visual features to be masked includes:

[0095] Based on the cross-attention mechanism, fuse the visual features and the language features to obtain the fused features, acquire the sequence of the fused features, and randomly select a position in the sequence;

[0096] Obtain the attention score map, and based on the attention score map, use the visual features before the position as the visual features to be masked.

[0097] For ease of explanation, an example is given below:

[0098] For visual feature V text and language feature L text , V text comes from visual feature extraction at the image level, while L text is obtained by performing feature extraction on the text sequence. To perform feature fusion between the visual and text modalities, the cross-attention mechanism is used for processing and fusion, enabling the model to focus on the relevant features in the two different modalities and learn how to match image regions with the relevant words in the descriptive text, thereby effectively combining these features to complete the text recognition task.

[0099] To perform the cross-attention mechanism, project the visual features and the language features into corresponding dimensions respectively, and generate three vectors of Query, Key, and Value by constructing learnable weight matrices W q 、W k 、W v .

[0100] Among them, the Chinese of Query, Key, and Value can be understood as query, key, and value respectively.

[0101] Among them, the text is used as the query, and the visual features are used as the key and the value.

[0102] Then, by calculating the product results of each feature vector in the text and the visual feature vectors in the image, and using the softmax function to normalize the product results, the attention score V is obtained. score , after obtaining the weights of the attention, multiply V score by the value value to obtain the final attention result V cross , as shown in the formula.

[0103] V cross = Cross(V text , L text ).

[0104] Among them, the visual feature is V text , the language feature is L text , and V cross is the attention result.

[0105] S204. Classify the masked visual features to generate a character sequence, map the character sequence to a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector;

[0106] Among them, the pre-trained model includes one or a combination of a Word2Vec model and a BRET model.

[0107] Among them, the step of classifying the masked visual features to generate a character sequence, mapping the character sequence to a text vector through the text encoder of the pre-trained model, and obtaining similar words generated by the pre-trained model based on the text vector includes:

[0108] Replace the feature values of the visual features to be masked with selected tokens to obtain the masked visual features;

[0109] Among them, a random probability is used to obtain the indexes of the visual features to be masked, so as to replace them with corresponding tokens. For example, if the ninth position is selected in the character sequence, then the top k visual features will be found in the order of the attention scores from high to low at the ninth position, and then the visual feature indexes will be obtained through random probability, so as to perform the replacement of the visual features.

[0110] This method randomly occludes a part of the image features of the image, enabling the model to learn more robust feature representations, and being able to extract key image information in the case of a noisy background, thereby improving the anti-interference and generalization capabilities of the model. Since the masking strategy is only performed in the training stage, it will not bring additional time overhead to the test and inference stages of the silk screen information recognition model.

[0111] Feed the masked visual features into a linear layer, and classify the masked visual features through the classification function to generate a character sequence corresponding to the masked visual features;

[0112] Map the character sequence into a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector.

[0113] For ease of explanation, the following is an example:

[0114] Feed the masked visual feature V mask into a linear layer, and classify it through softmax to generate corresponding character probability scores. By selecting the character with the highest probability score, combine to generate a predicted character sequence seq t ;

[0115] Use the pre-trained Word2Vec model to map the predicted character sequence seq t into a text vector seq c , capture the semantic similarity between characters, and generate similar words S for multiple characters through the BRET encoder t for sequence expansion;

[0116] For the similar words S1, S2, S3, S4... S of multiple character sequences t , perform text vector encoding through the CLIP text encoder to obtain corresponding text features C1, C2, C3, C4... C t ;

[0117] S205, optimize the screen printing information recognition model through the CLIP text encoder and the similar words;

[0118] Among them, the CLIP text encoder is the Contrastive Language-Image Pre-Training text encoder in full name.

[0119] Among them, optimizing the screen printing information recognition model through the CLIP text encoder and the similar words includes:

[0120] Encode the similar words through the CLIP text encoder to obtain the text features corresponding to the similar words. Use the cosine similarity formula to calculate the similarity between the visual features and the text features, and construct a similarity matrix through multiple similarities;

[0121] In the similarity matrix, regard the visual features and the text features with similarity higher than the preset value as positive sample pairs, and regard the visual features and the text features with similarity not higher than the preset value as negative sample pairs;

[0122] Using a contrastive loss function, optimize the silk screen information recognition model with the goal of minimizing the distance between the positive sample pairs and maximizing the distance between the negative sample pairs.

[0123] For ease of explanation, an example is as follows:

[0124] For each input batch n, in order to further align the visual feature V mask and the text feature seq c Calculate the similarity between all visual features and text features of this batch through the cosine of vectors, thereby generating an n×n similarity matrix s.

[0125] The calculation model of the similarity is as follows:

[0126]

[0127] where sim(v i , s j ) represents the similarity between v i , s j , v i represents the visual feature of the i-th chip image, v j represents the visual feature of the j-th chip image, s i and s j represent the text features corresponding to the image features v i and v j respectively.

[0128] Construct positive and negative sample pairs corresponding to the image by generating the similarity matrix s, where each element s ij represents the cosine similarity between the visual feature vector v i and the text feature vector s j , and the elements on the diagonal represent positive sample pairs, which are correctly matched image and text feature pairs;

[0129] The elements off the diagonal represent negative sample pairs, which are incorrectly matched image and text feature pairs.

[0130] v i represents the visual feature of the i-th chip image, v j represents the visual feature of the j-th chip image, s i and s j represent the text features corresponding to the image features v i and v j respectively. The method for constructing positive and negative sample pairs is as follows:

[0131]

[0132] Through this construction method, the time and cost of manual annotation can be reduced, and the generalization ability of the model can be improved.

[0133] Among them, a contrastive loss function is used to calculate the similarity degree of positive and negative sample pairs. By maximizing the similarity between similar features and minimizing the similarity between different features, the image visual features and text features of positive samples are closer in the common embedding space, while the image visual features and text features of negative samples are farther apart, thus effectively optimizing the learning ability of the silk screen information recognition model.

[0134] S206. Use the optimized silk screen information recognition model to obtain the predicted silk screen information, and train the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information.

[0135] Exemplarily, using the optimized silk screen information recognition model to obtain the predicted silk screen information, and training the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information includes:

[0136] Use the optimized silk screen information recognition model to obtain the predicted silk screen information corresponding to the masked visual features. Through the cross-entropy loss function, obtain the cross-entropy loss value between the predicted silk screen information and the true silk screen information, and train the silk screen information recognition model based on the cross-entropy loss value.

[0137] Among them, the corresponding relationship between the preset chip image and the true silk screen information is one-to-one, and different preset chip images correspond to different true silk screen information.

[0138] Among them, the true silk screen information is the correct silk screen information.

[0139] S207. When the cross-entropy loss value is less than the preset loss value, obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0140] Exemplarily, when the cross-entropy loss value is less than the preset loss value, obtaining the silk screen information recognition result output by the trained silk screen information recognition model includes:

[0141] When the cross-entropy loss value is less than the preset loss value, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0142] Among them, after obtaining the silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than the preset loss value, the chip silk screen information recognition method includes:

[0143] Create a display window and display the current text description information in the display window.

[0144] Compared with the existing methods for identifying chip silk screen information, the advantages of this application are as follows:

[0145] First, this application extracts the visual features and language features of the image through contrastive learning and pre-training of the silk screen information recognition model, reducing the manual annotation cost and time investment, better capturing the details and features of the chip silk screen information, and improving the overall recognition performance.

[0146] Second, through feature fusion and masking mechanism, this application can effectively establish associations between different modal features, train the ability of the silk screen information recognition model to perform text recognition in limited visual features, and improve the recognition accuracy of the silk screen information recognition model in complex and low-resolution scenarios.

[0147] Finally, this application utilizes the prior knowledge of the language silk screen information recognition model. By expanding similar words for the predicted character sequence, multiple possible candidates are provided for each predicted character sequence. The most appropriate result is selected according to the matching degree between the text language features of these candidates and the visual features of the image to guide the generation of the correct recognition result, improving the recognition accuracy and generalization ability of the silk screen information recognition model in various complex scenarios.

[0148] Please refer to Figure 3 , Figure 3 which is the flowchart for obtaining the current text description information provided by the embodiment of this application, and is described in detail as follows:

[0149] S301, when the cross-entropy loss value is less than the preset loss value, stop training the silk screen information recognition model and save the trained silk screen information recognition model;

[0150] S302, obtain the recognition accuracy of the trained silk screen information recognition model on the validation set;

[0151] S303, when the recognition accuracy is greater than the preset accuracy, obtain the current chip image and the current text description information corresponding to the current chip image, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0152] Among them, the silk screen information recognition effect of the model is evaluated on the validation set to ensure that it performs well in extracting image features, and according to the evaluation results, the model parameters and training strategies are dynamically adjusted.

[0153] For ease of explanation, the following is an example:

[0154] Search from high to low among a set of candidate values τ ∈ {0.05, 0.1, 0.2, 0.5, 0.75, 1.0}. As the training progresses, dynamically adjust τ according to the change of loss, and select an appropriate temperature coefficient τ through the performance metrics of the validation set.

[0155] During the training process of the model, adopt a step decay strategy. After every fixed number of training times, reduce the learning rate by a fixed ratio. This method can smooth the training process and avoid premature convergence and oscillation phenomena.

[0156]

[0157] Among them, η0 represents the initial learning rate, d is the decay factor, t represents the decay ratio, s represents the decay step, and η t represents the current learning rate.

[0158] Among them, 0 < d < 1. As t / s increases, the current learning rate will gradually decrease.

[0159] Among them, the decay factor determines the decay speed of the initial learning rate.

[0160] In the embodiment of the present application, when the recognition accuracy rate is greater than the preset accuracy rate, obtain the current chip image and the current text description information corresponding to the current chip image, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model. Since the preset accuracy rate is based on a custom function, the trained silk screen information recognition model can meet the recognition needs of users.

[0161] Corresponding to the chip silk screen information recognition method described in the above embodiment, please refer to Figure 4 , Figure 4 which is a schematic block diagram of the chip silk screen information recognition device provided by the embodiment of the present application. Figure 4 The chip silk screen information recognition device 400 shown can be applied to an electronic device in the application scenario diagram shown in Figure 1 . Taking the electronic device as an example, the chip silk screen information recognition device 400 shown in Figure 4 will be elaborated in detail. The chip silk screen information recognition device 400 may include a first acquisition module 401, an extraction module 402, a selection module 403, a second acquisition module 404, an optimization module 405, a third acquisition module 406, and an identification module 407.

[0162] The first acquisition module 401 is used to acquire a preset chip image and the preset text description information corresponding to the preset chip image;

[0163] An extraction module 402, configured to use the ResNet model and the Transformer model to extract features from the preset chip image to obtain visual features corresponding to the preset chip image, and use the Transformer model to extract features from the preset text description information to obtain language features corresponding to the preset chip image;

[0164] A selection module 403, configured to fuse the visual features and the language features based on a cross-attention mechanism to obtain fused features, obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked;

[0165] A second acquisition module 404, configured to classify the masked visual features to generate a character sequence, map the character sequence to a text vector through a text encoder of a pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector;

[0166] An optimization module 405, configured to optimize the silk screen information recognition model through a CLIP text encoder and the similar words;

[0167] A third acquisition module 406, configured to use the optimized silk screen information recognition model to obtain predicted silk screen information, and train the silk screen information recognition model according to a cross-entropy loss value between the predicted silk screen information and the true silk screen information;

[0168] A recognition module 407, configured to obtain a silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than a preset loss value.

[0169] It should be noted that the embodiments in this specification are all described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.

[0170] The beneficial effects of the embodiments of the present application are in two aspects. On the one hand, when the cross-entropy loss value is less than the preset loss value, a silk screen information recognition result output by the trained silk screen information recognition model is obtained. Since there is no need to manually define a template, the recognition time of the chip silk screen information is reduced, which is beneficial to improving the recognition efficiency of the chip silk screen information. On the other hand, through the feature masking technology, the visual features are masked, and the ability of the silk screen information recognition model to recognize silk screen information from limited visual features is trained. And by expanding the similar words, multiple possible candidates are provided for each predicted character sequence, improving the prediction ability of the silk screen information recognition model in a low-resolution scenario, and further improving the recognition accuracy of the silk screen information recognition model for the chip silk screen information.

[0171] Please refer to Figure 5 , Figure 5 , which is a schematic structural diagram of the electronic device provided by the embodiment of the present application.

[0172] As Figure 5 shown Figure 5 The electronic device 2 includes: at least one processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the at least one processor 20. When the processor 20 executes the computer program 22, the steps in any of the above method embodiments are implemented.

[0173] The electronic device 2 may include, but is not limited to, the processor 20 and the memory 21. Those skilled in the art can understand that Figure 5 merely an example of the electronic device 2, which does not constitute a limitation on the electronic device 2, and may include more or fewer components than shown in the figure, or combine some components, or different components. For example, it may also include input / output devices, network access devices, etc.

[0174] Among them, the processor 20 is used to run the computer program 22 stored in the memory 21, and when executing the computer program 22, the following steps are implemented:

[0175] Obtain a preset chip image and preset text description information corresponding to the preset chip image;

[0176] Use the ResNet model and the Transformer model to extract features from the preset chip image to obtain visual features corresponding to the preset chip image, and use the Transformer model to extract features from the preset text description information to obtain language features corresponding to the preset chip image;

[0177] Based on the cross-attention mechanism, fuse the visual features and the language features to obtain fused features, obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked;

[0178] Classify the masked visual features to generate a character sequence, map the character sequence to a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector;

[0179] Optimize the silk screen information recognition model through the CLIP text encoder and the similar words; use the optimized silk screen information recognition model to obtain the predicted silk screen information, and train the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information; when the cross-entropy loss value is less than the preset loss value, obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0180] In some embodiments, the processor 20 is used to implement:

[0181] Based on the cross-attention mechanism, fuse the visual feature and the language feature to obtain the fused feature, obtain the sequence of the fused feature, and randomly select a position in the sequence;

[0182] Obtain the attention score map, and according to the attention score map, use the visual feature before the position as the visual feature to be masked.

[0183] In some embodiments, the processor 20 is used to implement:

[0184] Replace the feature value of the visual feature to be masked with the selected token to obtain the masked visual feature;

[0185] Send the masked visual feature into the linear layer, and classify the masked visual feature through the classification function to generate the character sequence corresponding to the masked visual feature;

[0186] Map the character sequence into a text vector through the text encoder of the pre-trained model, and obtain the similar words generated by the pre-trained model based on the text vector.

[0187] In some embodiments, the processor 20 is used to implement:

[0188] Encode the similar words through the CLIP text encoder to obtain the text features corresponding to the similar words, use the cosine similarity formula to calculate the similarity between the visual feature and the text feature, and construct a similarity matrix through multiple similarities;

[0189] In the similarity matrix, regard the visual feature and the text feature with a similarity higher than the preset value as a positive sample pair, and regard the visual feature and the text feature with a similarity not higher than the preset value as a negative sample pair;

[0190] Use the contrastive loss function to optimize the silk screen information recognition model with the goal of minimizing the distance between the positive sample pairs and maximizing the distance between the negative sample pairs.

[0191] In some embodiments, the processor 20 is configured to:

[0192] When the cross-entropy loss value is less than the preset loss value, stop training the silk screen information recognition model and save the trained silk screen information recognition model;

[0193] Obtain the recognition accuracy of the trained silk screen information recognition model on the validation set;

[0194] When the recognition accuracy is greater than the preset accuracy, obtain the current chip image and the current text description information corresponding to the current chip image, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model.

[0195] In some embodiments, the processor 20 is configured to:

[0196] Create a display window and display the current text description information in the display window.

[0197] In some embodiments, the processor 20 is configured to:

[0198] The pre-trained model includes one or a combination of a Word2Vec model and a BRET model.

[0199] The so-called processor 20 may be a central processing unit (CPU), and this processor 20 may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or this processor may also be any conventional processor, etc.

[0200] The memory 21 may be an internal storage unit of the electronic device 2 in some embodiments, such as a hard disk or memory of the electronic device 2. The memory 21 may also be an external storage device of the electronic device 2 in other embodiments, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the electronic device 2. Further, the memory 21 may also include both the internal storage unit and the external storage device of the electronic device 2. The memory 21 is used to store an operating system, application programs, a Boot Loader, data, and other programs, such as the program code of the computer program. The memory 21 may also be used to temporarily store the data that has been output or will be output.

[0201] It should be noted that, regarding the information interaction, execution process, etc. between the above-mentioned device / unit, since it is based on the same concept as the method embodiment of the present application, for its specific functions and the technical effects brought, please refer to the method embodiment part specifically, and details will not be elaborated here.

[0202] The embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps in the above-mentioned various method embodiments can be implemented.

[0203] The program code is stored in the computer-readable storage medium, and the program code can be called by the processor to execute the chip silk screen information recognition method described in the above method embodiment.

[0204] The computer-readable storage medium has a storage space for the program code.

[0205] The program code includes the code of any step in the chip silk screen information recognition method described in the above method embodiment.

[0206] For the specific implementation of the above operations, please refer to the previous embodiments, and details will not be elaborated here.

[0207] Among them, the computer-readable storage medium may also be an external storage device of the chip silk screen information recognition device or the electronic device. For example, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, a non-transitory computer-readable storage medium, etc. equipped on the chip silk screen information recognition device or the electronic device.

[0208] Due to the computer program stored in the computer-readable storage medium, any chip silk screen information recognition method provided by the embodiments of the present application can be executed. Therefore, the computer-readable storage medium can achieve the beneficial effects that any chip silk screen information recognition method provided by the embodiments of the present application can achieve. For details, refer to the previous embodiments and will not be repeated here.

[0209] The embodiments of the present application provide a computer program product. When the computer program product runs on an electronic device, it causes the electronic device to execute the above-mentioned chip silk screen information recognition method.

[0210] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0211] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in a processing unit, or each unit exists physically alone, or two or more units are integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0212] Based on such an understanding, to implement all or part of the processes in the above-mentioned method embodiments of the present application, it can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can at least include: any entity or device that can carry the computer program code to an electronic device, a recording medium, a computer memory, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), an electrical carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disc, etc.

[0213] In the above embodiments, the descriptions of the various embodiments have their own focuses. For parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0214] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are similarly included in the patent protection scope of the present application.

Claims

1. A method for identifying chip screen printing information, characterized in that, Applied to an electronic device, the electronic device stores a silk screen information recognition model, the silk screen information recognition model includes a ResNet model and a Transformer model, and the chip silk screen information recognition method includes: Obtain a preset chip image and preset text description information corresponding to the preset chip image; Using the ResNet model and the Transformer model, perform feature extraction on the preset chip image to obtain visual features corresponding to the preset chip image, and using the Transformer model, perform feature extraction on the preset text description information to obtain language features corresponding to the preset chip image; Based on the cross-attention mechanism, fuse the visual features and the language features to obtain fused features, obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked; Classify the masked visual features to generate a character sequence, map the character sequence to a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector; Optimize the silk screen information recognition model through the CLIP text encoder and the similar words; Use the optimized silk screen information recognition model to obtain predicted silk screen information, and train the silk screen information recognition model according to the cross-entropy loss value between the predicted silk screen information and the true silk screen information; When the cross-entropy loss value is less than the preset loss value, obtain the silk screen information recognition result output by the trained silk screen information recognition model; Obtain a sequence of the fused features, randomly select a position in the sequence, and use the visual features before the position as the visual features to be masked, including: Obtain an attention score map, and based on the attention score map, use the visual features before the position as the visual features to be masked; Replace the feature values of the visual features to be masked with selected tokens to obtain the masked visual features; Among them, a random probability is used to obtain the index of the visual features to be masked, so as to replace them with corresponding tokens.

2. The chip silk-screen information recognition method according to claim 1, wherein The classifying the masked visual features to generate a character sequence, mapping the character sequence to a text vector through the text encoder of the pre-trained model, and obtaining similar words generated by the pre-trained model based on the text vector includes: Send the masked visual features into a linear layer, and classify the masked visual features through the classification function to generate a character sequence corresponding to the masked visual features; Map the character sequence to a text vector through the text encoder of the pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector.

3. The method for identifying chip silk-screen information according to claim 1, wherein The optimizing the silk screen information recognition model through the CLIP text encoder and the similar words includes: Encode the similar words through the CLIP text encoder to obtain the text features corresponding to the similar words, use the cosine similarity formula to calculate the similarity between the visual features and the text features, and construct a similarity matrix through multiple such similarities; In the similarity matrix, regard the visual features and the text features with similarities higher than the preset value as positive sample pairs, and regard the visual features and the text features with similarities not higher than the preset value as negative sample pairs; Use the contrastive loss function to optimize the silk screen information recognition model with the goal of minimizing the distance between the positive sample pairs and maximizing the distance between the negative sample pairs.

4. The chip silk screen information recognition method according to claim 1, characterized in that, When the cross-entropy loss value is less than the preset loss value, obtain the silk screen information recognition result output by the trained silk screen information recognition model, including: When the cross-entropy loss value is less than the preset loss value, stop training the silk screen information recognition model and save the trained silk screen information recognition model; Obtain the recognition accuracy of the trained silk screen information recognition model on the validation set; When the recognition accuracy is greater than the preset accuracy, obtain the current chip image and the current text description information corresponding to the current chip image, input the current chip image and the current text description information into the trained silk screen information recognition model, and obtain the silk screen information recognition result output by the trained silk screen information recognition model.

5. The chip screen printing information recognition method according to claim 4, characterized in that After obtaining the silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than the preset loss value, the chip silk screen information recognition method includes: Create a display window and display the current text description information in the display window.

6. The chip silk screen information recognition method according to any one of claims 1 to 5, characterized in that The pre-trained model includes one or a combination of a Word2Vec model and a BRET model.

7. An information recognition device for chip screen printing, characterized in that Applied to an electronic device, the electronic device stores a silk screen information recognition model, and the silk screen information recognition model includes a ResNet model and a Transformer model, including: A first acquisition module for acquiring a preset chip image and the preset text description information corresponding to the preset chip image; An extraction module for using the ResNet model and the Transformer model to extract features from the preset chip image to obtain the visual features corresponding to the preset chip image, and using the Transformer model to extract features from the preset text description information to obtain the language features corresponding to the preset chip image; A selection module for fusing the visual features and the language features based on the cross-attention mechanism to obtain a fused feature, obtaining a sequence of the fused features, randomly selecting a position in the sequence, and using the visual features before the position as the visual features to be masked; A second acquisition module, configured to classify the masked visual features to generate a character sequence, map the character sequence to a text vector through a text encoder of a pre-trained model, and obtain similar words generated by the pre-trained model based on the text vector; An optimization module, configured to optimize the silk screen information recognition model through a CLIP text encoder and the similar words; A third acquisition module, configured to use the optimized silk screen information recognition model to obtain predicted silk screen information, and train the silk screen information recognition model according to a cross-entropy loss value between the predicted silk screen information and the true silk screen information; An identification module, configured to obtain a silk screen information recognition result output by the trained silk screen information recognition model when the cross-entropy loss value is less than a preset loss value; Obtaining a sequence of the fusion features, randomly selecting a position in the sequence, and using the visual features before the position as the visual features to be masked, including: Obtaining an attention score map, and using the visual features before the position as the visual features to be masked according to the attention score map; Replacing the feature values of the visual features to be masked with selected tokens to obtain the masked visual features; Among them, a random probability is used to obtain the index of the visual features to be masked, so as to replace them with corresponding tokens.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the chip silk screen information recognition method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the chip silk screen information recognition method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Natural scene text detection method and system based on attention mechanism feature fusion and enhancement

    CN114255456A

  • Mask-based interactive enhanced image text recognition method and system

    CN117710986A