Image recognition method based on OCR technology

By combining the feature fusion of image and audio data, a multimodal neural network is constructed, which solves the problem of misidentification of traditional image recognition methods under noisy or incomplete data, and achieves more accurate and robust image recognition.

CN120356223APending Publication Date: 2025-07-22BEIJING YUNCHENG FINANCIAL INFORMATION SERVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510375338.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Traditional image recognition methods rely on image data, making it difficult to fully understand image content, especially when there are noisy or incomplete data, errors are prone to occur.

Method used

By acquiring images and corresponding audio data, extracting their respective features and fusing them, building a multimodal neural network, and using the trained recognition model for analysis to output the recognition report.

Benefits of technology

It improves the accuracy and robustness of image recognition, reduces the possibility of misidentification, and achieves a more comprehensive understanding of image content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356223A_ABST
    Figure CN120356223A_ABST
Patent Text Reader

Abstract

The invention provides an image recognition method based on an OCR technology, and belongs to the technical field of image processing, and the method comprises the steps: obtaining a plurality of image data and audio data corresponding to the image data, and extracting the image features and original audio features of the image data; carrying out feature fusion on the image features and the original audio features, designing a multi-modal neural network based on a feature fusion result, and constructing a recognition model according to the multi-modal neural network; and using the identification model to identify the to-be-identified image, and outputting an identification report, thereby improving the accuracy and robustness of identification, more comprehensively understanding the image content, reducing the possibility of misidentification, and improving the rationality of the identification result and the universality of the identification process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image recognition method based on OCR technology. Background Art

[0002] Image recognition technology is widely applied in multiple fields, including autonomous driving, medical image analysis, security monitoring, face recognition, image search, etc. Traditional image recognition methods usually rely on image data, but relying solely on image data may not fully understand the image content, and it is prone to errors when facing noise or incomplete data.

[0003] Therefore, the present invention provides an image recognition method based on OCR technology. Summary of the Invention

[0004] The image recognition method based on OCR technology provided by the present invention obtains an image and corresponding audio data, extracts their respective features, fuses the image features and audio features, constructs a multimodal neural network to enhance the information expression ability, uses the trained recognition model to analyze the image to be recognized, and outputs an accurate recognition report, thereby improving the accuracy and robustness of recognition, more comprehensively understanding the image content, reducing the possibility of misrecognition, and improving the rationality of the recognition result and the universality of the recognition process.

[0005] The present invention provides an image recognition method based on OCR technology, including: Step 1: Obtain a plurality of image data and audio data corresponding to the image data, and extract the image features and original audio features of the image data; Step 2: Perform feature fusion on the image features and the original audio features, design a multimodal neural network based on the feature fusion result, and then construct a recognition model according to the multimodal neural network; Step 3: Use the recognition model to recognize the image to be recognized and output a recognition report.

[0006] The present invention provides an image recognition method based on OCR technology, which obtains a plurality of image data and audio data corresponding to the image data, and extracts the image features and original audio features of the image data, including: Perform a first preprocessing on the obtained image data, perform a second preprocessing on the audio data, and add a connection identifier to the second preprocessing result based on the first preprocessing result; Perform a first classification of the image according to the first preprocessing result, and perform a second classification of the image based on the second preprocessing result; Further refine the first classification result based on the connection mark and the second classification result to obtain the actual classification of the image data; Extract the image features and original audio features of the image data based on the actual classification result.

[0007] The present invention provides an image recognition method based on OCR technology, which performs feature fusion on image features and original audio features, designs a multi-modal neural network based on the feature fusion result, and further constructs a recognition model according to the multi-modal neural network, including: Synthesize an image feature vector according to the image features, calculate a first feature dimension, synthesize an original audio feature vector according to the original audio features, and calculate a second feature dimension; Perform feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result; Determine the neural network architecture according to the feature fusion result, further determine the intermediate layer and the decision layer based on the neural network architecture, and construct a multi-modal neural network according to the neural network architecture, the intermediate layer and the decision layer, and then construct a recognition model.

[0008] The present invention provides an image recognition method based on OCR technology, which performs feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result, including: Based on the first feature dimension and the second feature dimension, use a synthesizer to generate a synthetic audio feature corresponding to the image features, compare the synthetic audio feature with the original audio feature for similarity, and obtain a similarity comparison result; Determine the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value; When the feature loss value is less than a preset threshold, perform a first fusion on the synthetic audio feature and the original audio feature. When the feature loss value is greater than the preset threshold, use the feature loss value to adjust and optimize the synthesizer parameters; Perform a second fusion on the first fusion result and the image features based on the first feature dimension and the second feature dimension to obtain a feature fusion result.

[0009] The present invention provides an image recognition method based on OCR technology, which determines the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value, including: , where represents the feature loss function; represents the important coefficient that controls the importance of the cross-entropy loss in the total loss; represents the original audio label corresponding to the i-th sample; represents the synthetic audio label corresponding to the i-th sample; N represents the total number of samples; i represents the sample index; represents the important coefficient that controls the importance of the cosine similarity loss in the total loss; A represents the synthetic audio feature vector; B represents the original audio feature vector; Represents the important coefficient of the control regularization term in the total loss; Represents the loss term of L2 regularization; Represents the square of the value of the j-th synthesizer parameter; j represents the index of the synthesizer parameter; M represents the total number of synthesizer parameters; Represents the important coefficient of the control mean squared error loss in the total loss; Represents the image feature vector corresponding to the i-th sample; Represents the important coefficient of the control multimodal fusion loss in the total loss; Represents the norm of the synthesized audio feature vector; Represents the norm of the original audio feature vector.

[0010] The present invention provides an image recognition method based on OCR technology, determines a neural network architecture according to the feature fusion result, further determines an intermediate layer and a decision layer based on the neural network architecture, constructs a multimodal neural network according to the neural network architecture, intermediate layer and decision layer, and then constructs a recognition model, including: Performs fused text feature extraction on the feature fusion result, conducts type analysis based on the extracted fused text features, and selects a neural network architecture from the type-architecture table according to the type analysis result; Determines the task complexity from the requirement-complexity table according to the user requirement and the type analysis result, determines the number of intermediate layers and the number of neurons corresponding to each layer according to the task complexity, and then obtains the intermediate layer; Determines the OCR target according to the user requirement, determines the corresponding processing function from the target-function table according to the OCR target, and then obtains the decision layer; Combines the selected neural network architecture with the intermediate layer and the decision layer to form a complete multimodal neural network, and then constructs a recognition model.

[0011] The present invention provides an image recognition method based on OCR technology, determines the OCR target according to the user requirement, determines the corresponding processing function from the target-function table according to the OCR target, and then obtains the decision layer, including: Clarifies the specific requirements of the OCR task according to the user requirement, and then obtains the OCR target, and selects a processing function from the target-function table according to the OCR target; Performs a first-level extraction on the type analysis result according to the processing function to obtain low-level fused text features, combines the processing function with the low-level fused text features to perform a second-level extraction on the type analysis result to obtain middle-level fused text features, and combines the processing function with the middle-level fused text features to perform a third-level extraction on the type analysis result to obtain high-level fused text features; Constructs a decision layer based on the high-level fused text features.

[0012] The present invention provides an image recognition method based on OCR technology, which uses the recognition model to recognize the image to be recognized and outputs a recognition report, including: Input the image to be recognized into the recognition model, determine the final recognition result according to the class label and model confidence output by the recognition model, and generate and output a recognition report according to the recognition result.

[0013] Compared with the prior art, the beneficial effects of the present application are as follows: By obtaining the image and the corresponding audio data, extracting their respective features, fusing the image features and the audio features, constructing a multi-modal neural network to enhance the information expression ability, using the trained recognition model to analyze the image to be recognized, and outputting an accurate recognition report, thereby improving the accuracy and robustness of the recognition, more comprehensively understanding the image content, reducing the possibility of misrecognition, and improving the rationality of the recognition result and the universality of the recognition process.

[0014] Other features and advantages of the present invention will be described in the following specification, and part of them will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention can be achieved and obtained by the structures specifically pointed out in the written specification and the drawings.

[0015] The technical solution of the present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings: Figure 1 is a schematic flowchart of an image recognition method based on OCR technology provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0018] Embodiment 1: The embodiment of the present invention provides an image recognition method based on OCR technology, as Figure 1 shown, including: Step 1: Obtain a plurality of image data and the audio data corresponding to the image data, and extract the image features and the original audio features of the image data; Step 2: Perform feature fusion on the image features and the original audio features, design a multi-modal neural network based on the feature fusion result, and further construct a recognition model according to the multi-modal neural network; Step 3: Use the recognition model to recognize the image to be recognized and output a recognition report.

[0019] In this embodiment, the image features are to extract the basic fusion text features, color histograms, shape features, etc. of the image. For example, extract the font features and positions of the text in the image. The original audio features are to extract the spectral features, time-domain features, etc. of the audio signal. For example, extract features such as the pitch and volume changes of the audio.

[0020] In this embodiment, the feature fusion result is to fuse the image feature vector and the audio feature vector. Usually, simple concatenation or more complex fusion methods (such as weighted average) can be used. For example, concatenate an image feature vector of length 128 with an audio feature vector of length 12 to obtain a feature fusion result of length 140.

[0021] In this embodiment, the multi-modal neural network is a neural network model that combines image and audio features. For example, construct a multi-layer perceptron, where the input layer receives the feature fusion result, and after being processed by the intermediate layer, the output layer gives the classification result.

[0022] In this embodiment, the synthesizer generates synthetic audio features corresponding to the image features, compares them with the original audio features to obtain a comparison result, determines the loss function of the synthesizer according to the comparison result to obtain a feature loss value. If the loss value is less than the threshold, perform the first fusion of the audio features; if it is greater than the threshold, optimize the synthesizer parameters, and perform the second fusion of the first fusion result and the image features to generate the final feature fusion result.

[0023] In this embodiment, the input of the recognition model is the feature fusion result (a vector of length 140), and the output is the classification result (such as "welcome" or "goodbye"). For example, the input is the feature fusion vector, and the output is the corresponding class label.

[0024] In this embodiment, perform fusion text feature extraction on the feature fusion result, and perform type analysis based on the extracted fusion text features. According to the type analysis result, select a suitable neural network architecture from the type-architecture table, determine the task complexity according to the user requirements and the type analysis result, and thus determine the number of intermediate layers and the number of neurons in each layer. Finally, select a processing function from the target-function table according to the OCR target to form a decision layer, and combine these layers to construct a complete multi-modal neural network recognition model.

[0025] In this embodiment, the image to be recognized refers to the input image that needs to be text-recognized by the OCR system, usually an image of a scanned document, a photo, or handwritten text. For example, a scanned document containing printed text with the content "Welcome to use OCR technology".

[0026] In this embodiment, the recognition report is a summary document generated by the model after recognizing the input image, containing information such as recognition results, category labels, and model confidence, and is usually used to display the recognition results to the user.

[0027] The working principle and beneficial effects of the above technical solution are as follows: By obtaining image and corresponding audio data, extracting their respective features, fusing the image features and audio features, constructing a multi-modal neural network to enhance the information expression ability, using the trained recognition model to analyze the image to be recognized, and outputting an accurate recognition report, so as to improve the accuracy and robustness of recognition, more comprehensively understand the image content, reduce the possibility of misrecognition, and improve the rationality of recognition results and the comprehensiveness of the recognition process.

[0028] Embodiment 2: The embodiment of the present invention provides an image recognition method based on OCR technology, which obtains a plurality of image data and audio data corresponding to the image data, and extracts the image features and original audio features of the image data, including: Perform a first preprocessing on the obtained image data, perform a second preprocessing on the audio data, and add a connection identifier to the second preprocessing result based on the first preprocessing result; Perform a first classification of the image according to the first preprocessing result, and perform a second classification of the image based on the second preprocessing result; Further refine the first classification result based on the connection marker and the second classification result to obtain the actual classification of the image data; Extract the image features and original audio features of the image data based on the actual classification result.

[0029] In this embodiment, the first preprocessing is to perform OCR-related preprocessing on the image data to improve the recognition accuracy, convert the color image to a grayscale image to reduce the computational complexity, use a filter to remove the noise in the image, convert the grayscale image to a black-and-white image to highlight the text area, and adjust the image size to meet the input requirements of the OCR model. For example, convert a color image containing text into a clear binary image for OCR recognition.

[0030] In this embodiment, the second preprocessing is to preprocess the audio data corresponding to the image to extract useful features, use noise cancellation technology to remove background noise, and extract features such as MFCC and spectrogram. For example, process the recorded audio file into an MFCC feature vector for subsequent analysis.

[0031] In this embodiment, adding a connection identifier is to generate a connection identifier for each pair of image and audio for subsequent classification and analysis. For example, generate an ID "img001_audio001" for the image and audio pair.

[0032] In this embodiment, the first classification of the image is a preliminary classification based on the preprocessed image data. OCR recognition is used to extract text. For example, the text recognized in the image is "Welcome"; the second classification of the image is a classification based on the preprocessed audio data. An acoustic model may be used. For example, the audio is classified as "human voice" or "background music".

[0033] In this embodiment, the classification result is further refined by using the connection identifier and the second classification result to further refine the first classification result to improve the classification accuracy. For example, if the image text is "Welcome" and the audio is "human voice", it is confirmed that the content of the image is a welcome message related to the human voice.

[0034] In this embodiment, the actual classification is obtained through a further refinement process to obtain the actual classification of the image data. For example, it is finally confirmed that the image content is "welcome slogan".

[0035] The working principle and beneficial effects of the above technical solution are as follows: The acquired image and audio data are preprocessed, feature extraction and classification are performed respectively, the preprocessing results are associated through a connection identifier, the image classification is further refined, and the image features and audio features are extracted based on the actual classification results, providing optimized input data for subsequent multi-modal fusion and recognition models, more accurately identifying the image data, and improving the adaptability of the overall model.

[0036] Embodiment 3: The embodiment of the present invention provides an image recognition method based on OCR technology, which performs feature fusion on image features and original audio features, designs a multi-modal neural network based on the feature fusion result, and then constructs a recognition model according to the multi-modal neural network, including: Synthesize an image feature vector according to the image features, calculate the first feature dimension, synthesize an original audio feature vector according to the original audio features, and calculate the second feature dimension; Perform feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result; Determine the neural network architecture according to the feature fusion result, further determine the intermediate layer and the decision layer based on the neural network architecture, and construct a multi-modal neural network according to the neural network architecture, the intermediate layer and the decision layer, and then construct a recognition model.

[0037] In this embodiment, after extracting the features of the image, the image feature vector is merged into a vector. Usually, a feature extraction algorithm (such as CNN) is used to obtain the features. For example, if the image features include a color histogram, texture features, and shape features, the finally synthesized image feature vector may be a vector with a length of 128, representing the combination of these features.

[0038] In this embodiment, the first feature dimension is the dimension for calculating the image feature vector. For example, if the length of the image feature vector is 128, then the first feature dimension is 128.

[0039] In this embodiment, the original audio feature vector is obtained by extracting audio features (such as MFCC, pitch, spectrum, etc.) and combining them into a single vector. For example, if the extracted audio features include 10 MFCC coefficients, pitch, and volume, the finally synthesized audio feature vector may be a vector with a length of 12.

[0040] In this embodiment, the second feature dimension is the dimension for calculating the audio feature vector. For example, if the length of the audio feature vector is 12, then the second feature dimension is 12.

[0041] In this embodiment, the neural network architecture is to design the overall structure of the network, including the input layer, intermediate layer, and decision layer. For example, a multi-layer perceptron (MLP) is selected as the neural network architecture, and the input layer receives the feature fusion result with a length of 140.

[0042] In this embodiment, the intermediate layer is designed with multiple hidden layers as needed, and the number of neurons in each layer can be adjusted according to the complexity of the features. For example, two hidden layers can be designed, with 64 neurons in the first layer and 32 neurons in the second layer.

[0043] In this embodiment, the decision layer is the output layer, which is usually used for classification or regression tasks. The number of output neurons corresponds to the number of task categories. For example, if the task is binary classification, the decision layer has 2 neurons and uses the softmax activation function.

[0044] The working principle and beneficial effects of the above technical solution are as follows: By synthesizing the image feature vector and the audio feature vector, calculating their respective feature dimensions, fusing the image and audio features based on these two dimensions to generate a feature fusion result, determining the neural network architecture according to the fusion result, designing the intermediate layer and the decision layer, and finally constructing a multi-modal neural network and an identification model to achieve more accurate identification tasks, which has a wide range of application scenarios.

[0045] Embodiment 4: The embodiment of the present invention provides an image recognition method based on OCR technology, which performs feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result, including: Based on the first feature dimension and the second feature dimension, use a synthesizer to generate a synthetic audio feature corresponding to the image feature, and perform a similarity comparison between the synthetic audio feature and the original audio feature to obtain a similarity comparison result; Determine the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value; When the feature loss value is less than the preset threshold, the synthetic audio feature and the original audio feature are fused for the first time. When the feature loss value is greater than the preset threshold, the feature loss value is used to adjust and optimize the synthesizer parameters; Based on the first feature dimension and the second feature dimension, the first fusion result and the image feature are fused for the second time to obtain the feature fusion result.

[0046] In this embodiment, the synthetic audio feature is generated based on the image feature and the audio feature dimension using a synthesizer (such as a generative adversarial network or a variational autoencoder) to generate an audio feature corresponding to the image feature. For example, if the image feature vector has a length of 128, the synthetic audio feature vector generated by the synthesizer may also have a length of 12, representing the audio feature related to the image content.

[0047] In this embodiment, the similarity comparison is to compare the similarity between the synthetic audio feature and the original audio feature. Usually, cosine similarity, Euclidean distance or other similarity metrics are used. For example, the cosine similarity between the synthetic audio feature and the original audio feature is calculated to obtain a value (ranging from -1 to 1), such as 0.85, indicating a high similarity between the two.

[0048] In this embodiment, the feature loss value is to calculate the loss function of the synthesizer according to the similarity comparison result and the image feature. Usually, mean squared error (MSE) or cross-entropy loss, etc. are used. For example, if the similarity is low, the loss function may be calculated as the MSE between the synthetic audio feature and the original audio feature, and the feature loss value is obtained as 0.05.

[0049] In this embodiment, the preset threshold is a standard for judging the feature loss value, usually set according to experiments. For example, the threshold is set to 0.1, indicating that if the feature loss value is less than 0.1, it means that the similarity between the synthetic audio feature and the original audio feature is high.

[0050] In this embodiment, the first fusion is to fuse the synthetic audio feature and the original audio feature when the feature loss value is less than the preset threshold, usually through simple concatenation or weighted average. For example, if the feature loss value is 0.05 (less than 0.1), the synthetic audio feature and the original audio feature are concatenated to obtain a new fused feature vector.

[0051] In this embodiment, the adjustment and optimization is to use the feature loss value to adjust and optimize the parameters of the synthesizer when the feature loss value is greater than the preset threshold, so as to improve the quality of the generated synthetic audio feature. For example, if the feature loss value is 0.15 (greater than 0.1), the parameters of the synthesizer are adjusted according to the backpropagation algorithm to reduce the loss.

[0052] In this embodiment, the second fusion is based on the first feature dimension and the second feature dimension, and the first fusion result and the image features are fused for the second time to obtain the final feature fusion result. For example, the first fusion result (including the synthesized audio features and the original audio features) is concatenated with the image feature vector to obtain the final feature fusion result, and the length may be 140 (128 + 12).

[0053] The working principle and beneficial effects of the above technical solution are as follows: The synthesizer generates synthesized audio features corresponding to the image features, compares them with the original audio features to obtain a comparison result, determines the loss function of the synthesizer according to the comparison result to obtain a feature loss value. If the loss value is less than the threshold, the first fusion of the audio features is performed; if it is greater than the threshold, the parameters of the synthesizer are optimized, and the first fusion result and the image features are fused for the second time to generate the final feature fusion result, improving the adaptability and performance of the model. The multiple fusion processes make the final feature fusion result more informative and improve the expressiveness of the model.

[0054] Embodiment 5: The embodiment of the present invention provides an image recognition method based on OCR technology, which determines the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value, including: , where represents the feature loss function; represents the important coefficient that controls the importance of the cross-entropy loss in the total loss; represents the original audio label corresponding to the i-th sample; represents the synthesized audio label corresponding to the i-th sample; N represents the total number of samples; i represents the sample index; represents the important coefficient that controls the importance of the cosine similarity loss in the total loss; A represents the synthesized audio feature vector; B represents the original audio feature vector; represents the important coefficient that controls the importance of the regularization term in the total loss; represents the loss term of L2 regularization; represents the square of the value of the j-th synthesizer parameter; j represents the index of the synthesizer parameter; M represents the total number of synthesizer parameters; represents the important coefficient that controls the importance of the mean square error loss in the total loss; represents the image feature vector corresponding to the i-th sample; represents the important coefficient that controls the importance of the multi-modal fusion loss in the total loss; represents the norm of the synthesized audio feature vector; represents the norm of the original audio feature vector.

[0055] In this embodiment, the original audio label and the synthesized audio label. For example, if the first sample is "Welcome", the original audio label = 1, indicating "Welcome", and the synthesized audio label may be the result predicted by the model, such as 0.9, indicating the confidence of the model in "Welcome".

[0056] The working principle and beneficial effects of the above technical solution are as follows: By defining a feature loss function, multiple loss terms are combined together, including cross-entropy loss, cosine similarity loss, L2 regularization loss, mean square error loss, and multi-modal fusion loss. Each loss term is weighted by a corresponding coefficient to ensure that each part occupies an appropriate proportion in the total loss. By optimizing this loss function and adjusting the synthesizer parameters, the similarity between the synthesized audio features and the original audio features is improved, the recognition ability of the model is enhanced, the synergy between different modal features is enhanced, and the robustness of the final recognition result is improved.

[0057] Embodiment 6: The embodiment of the present invention provides an image recognition method based on OCR technology. A neural network architecture is determined according to the feature fusion result, an intermediate layer and a decision layer are further determined based on the neural network architecture, and a multi-modal neural network is constructed according to the neural network architecture, the intermediate layer and the decision layer, and then an identification model is constructed, including: Perform fused text feature extraction on the feature fusion result, perform type analysis based on the extracted fused text features, and select a neural network architecture from the type-architecture table according to the type analysis result; Determine the task complexity from the requirement-complexity table according to the user requirements and the type analysis result, determine the number of intermediate layers and the number of neurons corresponding to each layer according to the task complexity, and then obtain the intermediate layer; Determine the OCR target according to the user requirements, determine the corresponding processing function from the target-function table according to the OCR target, and then obtain the decision layer; Combine the selected neural network architecture with the intermediate layer and the decision layer to form a complete multi-modal neural network, and then construct an identification model.

[0058] In this embodiment, the type analysis result is to classify the fused text features to identify their types (such as sentiment analysis, topic classification, etc.). For example, the analysis result shows that the fused text features belong to the "sentiment analysis" type.

[0059] In this embodiment, the type-architecture table is based on the mapping relationship between the text type and the suitable neural network architecture.

[0060] In this embodiment, user requirements refer to the expectations and requirements of users for the model, such as accuracy, speed, interpretability, etc. For example, users hope that the model can achieve an accuracy rate of over 85% in sentiment analysis tasks.

[0061] In this embodiment, the requirement-complexity table is mapped according to the relationship between user requirements and task complexity. For example, the table header is user requirements and task complexity. When the user requirement is high accuracy, the corresponding task complexity is high complexity.

[0062] Task complexity: The task difficulty determined according to type analysis and user requirements. For example, if the user requirement is high accuracy and the text type is sentiment analysis, the task complexity may be evaluated as "high".

[0063] In this embodiment, the middle layer refers to the number of hidden layers and the number of neurons in each layer in the neural network, which is usually related to task complexity. For example, for high-complexity tasks, 4 middle layers may be selected, with 256, 128, 64, and 32 neurons in each layer respectively.

[0064] In this embodiment, the decision layer is the output layer of the neural network, responsible for the final classification or regression task. For example, for sentiment analysis tasks, the decision layer may be a Softmax layer, which is used to output the probability of each sentiment (such as positive, neutral, negative).

[0065] In this embodiment, the OCR target is the specific target of optical character recognition, such as recognizing text, extracting key information, etc. For example, the OCR target may be to extract text information from an image and perform sentiment analysis.

[0066] In this embodiment, the target-function table is based on the mapping relationship between the OCR target and the corresponding processing function. For example, the table header is the OCR target and the processing function. When the OCR target is to extract text, the corresponding processing function is CRNN.

[0067] In this embodiment, the processing function is an algorithm or model used to achieve a specific OCR target. For example, if the OCR target is to extract text and perform sentiment analysis, LSTM is selected as the processing function.

[0068] The working principle and beneficial effects of the above technical solution are as follows: Fusion text features are extracted from the feature fusion result, and type analysis is performed based on the extracted fusion text features. According to the type analysis result, a suitable neural network architecture is selected from the type-architecture table. The task complexity is determined according to user requirements and type analysis results, thereby determining the number of middle layers and the number of neurons in each layer. The processing function is selected from the target-function table according to the OCR target to form the decision layer. These layers are combined to build a complete multi-modal neural network recognition model, which has good flexibility and can adapt to different recognition tasks.

[0069] Example 7: The embodiment of the present invention provides an image recognition method based on OCR technology. Determine the OCR target according to user needs, and determine the corresponding processing function from the target-function table, and then obtain the decision-making layer, including: Clarify the specific requirements of the OCR task according to user needs, and then obtain the OCR target. Select a processing function from the target-function table according to the OCR target; Perform the first-level extraction on the type analysis result according to the processing function to obtain the low-level fusion text feature. Combine the processing function with the low-level fusion text feature to perform the second-level extraction on the type analysis result to obtain the middle-level fusion text feature. Combine the processing function with the middle-level fusion text feature to perform the third-level extraction on the type analysis result to obtain the high-level fusion text feature; Construct the decision-making layer based on the high-level fusion text feature.

[0070] In this embodiment, the first-level extraction is to perform a preliminary extraction on the type analysis result using the selected processing function to obtain the low-level fusion text feature. For example, use the CRNN processing function to extract features from the original text to obtain the low-level fusion text feature represented as [0.1, 0.4, 0.3].

[0071] In this embodiment, the low-level fusion text feature is to perform a preliminary fusion of features from different sources. For example, combine image features and audio features with the low-level text features to obtain the low-level fusion feature [0.1, 0.4, 0.3, 0.2, 0.5, 0.1].

[0072] The second-level extraction is to perform a further extraction by combining the processing function with the low-level fusion text feature to obtain the middle-level fusion text feature. For example, use the LSTM processing function to process the low-level fusion text feature to obtain the middle-level fusion text feature represented as [0.3, 0.7, 0.5].

[0073] In this embodiment, the middle-level fusion text feature is to obtain a richer feature representation through further processing and combining more context information. For example, the middle-level fusion feature may be [0.3, 0.7, 0.5, 0.4, 0.6].

[0074] In this embodiment, the third-level extraction is to perform a final extraction by combining the processing function with the middle-level fusion text feature to obtain the high-level fusion text feature. For example, use the LSTM processing function again to process the middle-level fusion text feature to obtain the high-level fusion text feature represented as [0.5, 0.8, 0.6].

[0075] In this embodiment, the high-level fused text feature is the final feature representation, which contains multi-level information and context. For example, the high-level fused feature may be [0.5, 0.8, 0.6, 0.7], which is used for the subsequent decision-making layer.

[0076] The working principle and beneficial effects of the above technical solution are as follows: The specific objective of the OCR task is clarified according to the user's needs, and the corresponding processing function is selected from the objective-function table. The selected processing function is used to hierarchically extract the type analysis results to obtain low-level, middle-level, and high-level fused text features in sequence. The extraction of each level is based on the features of the previous level. Finally, the high-level fused text feature is used to construct the decision-making layer, so as to achieve effective recognition and classification of the OCR task, which helps to improve the information expression ability and recognition accuracy.

[0077] Embodiment 8: The embodiment of the present invention provides an image recognition method based on OCR technology, which uses the recognition model to recognize the image to be recognized and outputs a recognition report, including: Input the image to be recognized into the recognition model, determine the final recognition result according to the category label and model confidence output by the recognition model, and generate and output a recognition report according to the recognition result.

[0078] In this embodiment, the category label refers to the label generated after the model classifies the text content extracted from the image to be recognized, usually the semantic or emotional category of the text. For example, for the above image, the category label that the model may output is "positive" or "neutral", indicating the emotional tendency of the text.

[0079] In this embodiment, the model confidence refers to the confidence level of the model in its output category label, usually expressed in the form of a percentage or probability. The higher the confidence level, the more certain the model is about the judgment of this category. For example, if the model outputs a category label of "positive" and the confidence level is 85%, this means that the model believes that the probability that this text is of positive emotion is 85%.

[0080] The working principle and beneficial effects of the above technical solution are as follows: Input the image to be recognized into the multi-modal recognition model. The model analyzes the image features and audio features and outputs the corresponding category label and its confidence level. According to the output category label and confidence level, determine the final recognition result, and generate and output a detailed recognition report based on the recognition result, providing comprehensive information about the recognition process and result, and improving the recognition accuracy.

[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. An image recognition method based on OCR technology, characterized in that, Including: Step 1: Obtain multiple image data and the audio data corresponding to the image data, and extract the image features and the original audio features of the image data; Step 2: Perform feature fusion on the image features and the original audio features, design a multi-modal neural network based on the feature fusion result, and then construct an identification model according to the multi-modal neural network; Step 3: Use the identification model to identify the image to be identified and output an identification report.

2. The image recognition method based on OCR technology according to claim 1, wherein Obtaining multiple image data and the audio data corresponding to the image data, and extracting the image features and the original audio features of the image data, including: Perform a first preprocessing on the obtained image data, perform a second preprocessing on the audio data, and add a connection identifier to the second preprocessing result based on the first preprocessing result; Perform a first classification of the image according to the first preprocessing result, and perform a second classification of the image based on the second preprocessing result; Further refine the first classification result based on the connection marker and the second classification result to obtain the actual classification of the image data; Extract the image features and the original audio features of the image data based on the actual classification result.

3. The image recognition method based on OCR technology according to claim 1, characterized in that, Performing feature fusion on the image features and the original audio features, designing a multi-modal neural network based on the feature fusion result, and then constructing an identification model according to the multi-modal neural network, including: Synthesize an image feature vector according to the image features, calculate the first feature dimension, synthesize an original audio feature vector according to the original audio features, and calculate the second feature dimension; Perform feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result; Determine the neural network architecture according to the feature fusion result, further determine the intermediate layer and the decision layer based on the neural network architecture, construct a multi-modal neural network according to the neural network architecture, the intermediate layer and the decision layer, and then construct an identification model.

4. The image recognition method based on OCR technology according to claim 3, characterized in that Performing feature fusion on the image features and the original audio features based on the first feature dimension and the second feature dimension to obtain a feature fusion result, including: Based on the first feature dimension and the second feature dimension, use a synthesizer to generate a synthetic audio feature corresponding to the image features, compare the synthetic audio feature with the original audio features for similarity, and obtain a similarity comparison result; Determine the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value; When the feature loss value is less than a preset threshold, perform a first fusion on the synthetic audio feature and the original audio features. When the feature loss value is greater than the preset threshold, adjust and optimize the synthesizer parameters using the feature loss value; Perform a second fusion on the first fusion result and the image features based on the first feature dimension and the second feature dimension to obtain a feature fusion result.

5. The image recognition method based on OCR technology according to claim 4, characterized in that, Determine the loss function of the synthesizer according to the similarity comparison result in combination with the image features to obtain a feature loss value, Including: , where represents the feature loss function; represents the important coefficient that controls the importance of the cross-entropy loss in the total loss; represents the original audio label corresponding to the i-th sample; represents the synthesized audio label corresponding to the i-th sample; N represents the total number of samples; i represents the sample index; represents the important coefficient that controls the importance of the cosine similarity loss in the total loss; A represents the synthesized audio feature vector; B represents the original audio feature vector; represents the important coefficient that controls the importance of the regularization term in the total loss; represents the loss term of L2 regularization; represents the square of the value of the j-th synthesizer parameter; j represents the index of the synthesizer parameter; M represents the total number of synthesizer parameters; represents the important coefficient that controls the importance of the mean squared error loss in the total loss; represents the image feature vector corresponding to the i-th sample; represents the important coefficient that controls the importance of the multi-modal fusion loss in the total loss; represents the norm of the synthesized audio feature vector; represents the norm of the original audio feature vector.

6. The image recognition method based on OCR technology according to claim 3, wherein, Determine the neural network architecture according to the feature fusion result, further determine the intermediate layer and the decision layer based on the neural network architecture, construct a multi-modal neural network according to the neural network architecture, the intermediate layer and the decision layer, and then construct an identification model, including: Extract the fused text features from the feature fusion result, perform type analysis based on the extracted fused text features, and select a neural network architecture from the type-architecture table according to the type analysis result; Determine the task complexity from the requirement-complexity table according to the user requirements and the type analysis result, determine the number of intermediate layers and the number of neurons corresponding to each layer according to the task complexity, and then obtain the intermediate layers; Determine the OCR target according to the user requirements, determine the corresponding processing function from the target-function table according to the OCR target, and then obtain the decision layer; Combine the selected neural network architecture with the intermediate layer and the decision layer to form a complete multi-modal neural network, and then construct an identification model.

7. The image recognition method based on OCR technology according to claim 6, wherein Determine the OCR target according to the user requirements, determine the corresponding processing function from the target-function table according to the OCR target, and then obtain the decision layer, including: Clarify the specific requirements of the OCR task according to the user requirements, and then obtain the OCR target. Select a processing function from the target-function table according to the OCR target; Perform the first-level extraction on the type analysis result according to the processing function to obtain the low-level fused text features, combine the processing function with the low-level fused text features to perform the second-level extraction on the type analysis result to obtain the middle-level fused text features, and combine the processing function with the middle-level fused text features to perform the third-level extraction on the type analysis result to obtain the high-level fused text features; Construct the decision layer based on the high-level fused text features.

8. The image recognition method based on OCR technology according to claim 1, wherein Use the identification model to identify the image to be identified and output an identification report, including: Input the image to be identified into the identification model, determine the final identification result according to the class label and the model confidence output by the identification model, and generate and output an identification report according to the identification result.