Input image category identification method and device, computer equipment and storage medium

By combining OCR model and polygon fitting technology to extract text information in the image and using a hierarchical feature fusion method for recognition, the accuracy of image category recognition in the prior art is solved, and a more efficient recognition effect is achieved.

CN120125915APending Publication Date: 2025-06-10CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510319303.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, when facing complex and changeable image data, it is difficult to accurately identify image categories, resulting in misjudgment or misjudgment.

Method used

The pre-trained OCR model is used to combine polygon fitting technology for text positioning and segmentation, extract text information, and fuse visual features and text features through a hierarchical feature fusion method to perform similarity calculations for category division.

Benefits of technology

Accurate category recognition of the input image is realized, and the reliability of the recognition results is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125915A_ABST
    Figure CN120125915A_ABST
Patent Text Reader

Abstract

The invention relates to an input image category identification method and device, computer equipment and a storage medium, and the method comprises the following steps: carrying out the preprocessing of an input image, and obtaining a standard input image; extracting character information from the standard input image based on a pre-trained OCR model in combination with a polygon fitting technology; performing preliminary classification according to keywords, themes and context information of the text information to obtain a preliminary classification result; performing feature fusion on the visual features, the text features and the preliminary classification result of the standard input image by adopting a hierarchical feature fusion method to obtain image fusion features; and performing category division on the standard input image according to the image fusion feature and a pre-stored image template feature to obtain an input image category. The method and the device can be applied to financial service system application scenes, and can effectively realize accurate category identification of input images in a financial service system so as to facilitate subsequent system image information input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to a method, apparatus, computer device, and storage medium for recognizing the category of an input image. Background Art

[0002] With the rapid development of technology, the demand for image information processing is increasing day by day. Especially in core aspects such as policy images and medical record management, the system's ability to accurately classify and recognize input images is particularly crucial.

[0003] In a financial business system, complex image data (such as policy images including images and texts, transaction vouchers, etc.) may cause difficulties in system recognition due to factors such as lighting, angle, and occlusion. Similarly, in a health care system, due to the diversity of medical record and report images, for example, the complexity of pathological pictures and the inconsistencies of different organs and positions, it is difficult for the system to effectively classify image categories during the recognition process.

[0004] Although existing image recognition technologies and user process monitoring systems have made certain progress, there are still certain limitations when facing complex and variable image data. For example, the preset image feature matching rules and image analysis methods may not be able to comprehensively cover diverse image types and scenarios, resulting in possible misjudgments or missed judgments when the system classifies and judges input images in the system. Summary of the Invention

[0005] The purpose of the embodiments of this application is to propose a method, apparatus, computer device, and storage medium for recognizing the category of an input image to solve the problem of inability to accurately recognize the category of an input image.

[0006] In a first aspect, the embodiments of this application provide a method for recognizing the category of an input image, adopting the following technical solutions:

[0007] Obtain an input image, perform preprocessing on the input image to obtain a standard input image;

[0008] Based on a pre-trained OCR model combined with polygon fitting technology, perform text localization and segmentation on the standard input image, and extract text information;

[0009] Convert the text information into text data, and based on keywords, themes, and context information of the text data, perform a preliminary classification on the standard input image to obtain a preliminary classification result;

[0010] Extract the visual features of the standard input image and the text features of the text information, and use a hierarchical feature fusion method to fuse the visual features, the text features, and the preliminary classification result to obtain image fusion features;

[0011] Calculate the similarity between the image fusion features and the pre-stored image template features to obtain a similarity result, and classify the standard input image based on the similarity result to obtain the input image category.

[0012] In a second aspect, an input image category recognition device according to an embodiment of the present application further adopts the following technical solution:

[0013] An image processing module, configured to obtain an input image, preprocess the input image to obtain a standard input image;

[0014] A text extraction module, configured to perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with polygon fitting technology, and extract text information;

[0015] A preliminary classification module, configured to convert the text information into text information, and perform preliminary classification on the standard input image according to keywords, themes, and context information of the text information to obtain a preliminary classification result;

[0016] A feature fusion module, configured to extract the visual features of the standard input image and the text features of the text information, and use a hierarchical feature fusion method to fuse the visual features, the text features, and the preliminary classification result to obtain image fusion features;

[0017] A category division module, configured to calculate the similarity between the image fusion features and the pre-stored image template features to obtain a similarity result, and classify the standard input image based on the similarity result to obtain the input image category.

[0018] In a third aspect, an embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0019] A computer device includes a memory and a processor, and computer-readable instructions are stored in the memory. When the processor executes the computer-readable instructions, the steps of the input image category recognition method described in any one of the above are implemented.

[0020] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0021] A computer-readable storage medium stores computer-readable instructions thereon, and when the computer-readable instructions are executed by a processor, the steps of the input image category recognition method described in any one of the above are implemented.

[0022] Compared with the prior art, the embodiments of the present application mainly have the following beneficial effects: In this embodiment, by obtaining an input image, preprocessing the input image to obtain a standard input image; based on a pre-trained OCR model combined with a polygon fitting technique, performing text localization and segmentation on the standard input image to extract text information; converting the text information into text information, and based on the keywords, themes, and context information of the text information, performing a preliminary classification on the standard input image to obtain a preliminary classification result; extracting the visual features of the standard input image and the text features of the text information, and using a hierarchical feature fusion method to perform feature fusion on the visual features, the text features, and the preliminary classification result to obtain an image fusion feature; calculating the similarity between the image fusion feature and the pre-stored image template feature to obtain a similarity result, and based on the similarity result, performing category division on the standard input image to obtain the input image category. Thus, it effectively realizes accurate category recognition of the input image and improves the reliability of the recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the solutions in the present application, the following will briefly introduce the drawings required for the description of the embodiments of the present application. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is an exemplary system architecture diagram to which the present application can be applied;

[0025] Figure 2 A flowchart of an embodiment of the input image category recognition method according to the present application;

[0026] Figure 3 It is a schematic structural diagram of an embodiment of the input image category recognition device according to the present application;

[0027] Figure 4 It is a schematic structural diagram of an embodiment of the computer device according to the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this application belongs; the terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" and any variations thereof in the specification, claims, and drawings of this application are intended to cover non-exclusive inclusion. The terms "first", "second", etc. in the specification, claims, or drawings of this application are used to distinguish different objects and not to describe a specific order.

[0029] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of this application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an unrelated or alternative embodiment mutually exclusive of other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.

[0030] To enable those skilled in the art to better understand the solutions of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings.

[0031] As Figure 1 shown, the system architecture 100 may include a terminal device 101, a network 102, and a server 103. The terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0032] The user may use the terminal device 101 to interact with the server 103 through the network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.

[0033] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, the tablet computer 1012, or the mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop portable computer, a desktop computer, and the like.

[0034] The server 103 can be a server that provides various services, such as a background server that supports the pages displayed on the terminal device 101.

[0035] It should be noted that the input image category recognition method provided by the embodiments of the present application is generally executed by the server / terminal device. Correspondingly, the input image category recognition device is generally set in the server / terminal device.

[0036] It should be understood that Figure 1 the numbers of the terminal devices, the network, and the server in

[0037] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers. Figure 2 Continuing to refer to

[0038] shows a flowchart of an embodiment of the input image category recognition method according to the present application. The input image category recognition method includes the following steps:

[0039] In this embodiment, the input image refers to the image data sent or uploaded by the user to the business system platform. For example, in the online vehicle insurance system, the user will send or upload the vehicle insurance document images (including insurance policies, accident scene photos, vehicle damage photos, etc.) for the business agent to enter into the system. Or, in the medical and health platform, the user will upload medical images (including X-rays, CT images, medical record images, etc.) for the background to enter into the system. The preprocessing of the input image includes grayscale conversion, binarization, denoising processing, and image correction. By performing the preprocessing including the above steps on the input image, a standard and effective standard input image can be obtained.

[0040] Step S20, perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with polygon fitting technology, and extract text information;

[0041] In this embodiment, the standard input image contains text. The text in the standard input image is regionally located through a pre-trained OCR model to identify the text region in the standard input image, and then the identified text region is polygonally fitted to more accurately represent the shape of the text, so as to facilitate the segmentation and extraction of text information.

[0042] Step S30: Convert the text information into text data, and preliminarily classify the standard input image according to the keywords, themes, and context information of the text data to obtain a preliminary classification result.

[0043] In this embodiment, the text information corresponds to the image data of the text shape. The corresponding text content is extracted from the image data of the text shape through OCR technology to obtain text data. Keywords can be extracted from the text data through the BERT language model, and the context information can be understood. The theme can be extracted from the text data through the LDA topic modeling algorithm. According to the extracted keywords, themes, and context information, the standard input image is preliminarily classified according to predefined classification rules to obtain a preliminary classification result. The predefined classification rules can be formulated according to the actual situation. For example, if the text data contains "policy number" and "amount", it is classified as "policy"; if the text data contains "medical record number" and "diagnosis result", it is classified as "medical record".

[0044] Step S40: Extract the visual features of the standard input image and the text features of the text data, and use a hierarchical feature fusion method to fuse the visual features, the text features, and the preliminary classification result to obtain an image fusion feature.

[0045] In this embodiment, visual features are extracted from the standard input image through a preset visual encoder, and text features are extracted from the text data through a pre-trained language model (BERT). The text features include text content, keywords, phrases, text position information, semantic features, etc. Before the preliminary classification result is fused with features, it needs to be first converted into a result feature vector, and then feature fusion is performed according to the result feature vector.

[0046] Step S50: Calculate the similarity between the image fusion feature and the pre-stored image template feature to obtain a similarity result, and classify the standard input image based on the similarity result to obtain the input image category.

[0047] In this embodiment, the pre-stored image template features refer to the feature information formed by performing various feature processes on sample images. For example, in the online vehicle insurance system, by performing various feature processes on historical vehicle insurance policies, the image template features corresponding to the vehicle insurance policies are obtained. In the medical and health platform, by performing various feature processes on medical record images, the image template features corresponding to the medical record images are obtained. The similarity result is index data that measures the similarity degree between the image fusion features and the image template features, and the similarity between the objects corresponding to the above two features is reflected through the similarity result. The similarity result corresponds to a similarity value. The higher the similarity value, the more similar the image corresponding to the image fusion features and the image corresponding to the image template features are. Conversely, the less similar the image corresponding to the image fusion features and the image corresponding to the image template features are.

[0048] In this embodiment, by obtaining an input image, preprocessing the input image to obtain a standard input image; based on a pre-trained OCR model combined with polygon fitting technology, performing text localization and segmentation on the standard input image to extract text information; converting the text information into text data, and according to the keywords, themes, and context information of the text data, performing a preliminary classification on the standard input image to obtain a preliminary classification result; extracting the visual features of the standard input image and the text features of the text data, and using a hierarchical feature fusion method to perform feature fusion on the visual features, the text features, and the preliminary classification result to obtain image fusion features; calculating the similarity between the image fusion features and the pre-stored image template features to obtain a similarity result, and based on the similarity result, performing category division on the standard input image to obtain the input image category. Thus, the accurate category recognition of the input image is effectively realized, and the reliability of the recognition result is improved.

[0049] The method of this embodiment can be applied to the recognition of the input image category in the online vehicle insurance system. The user uploads the vehicle insurance document image (including images and text) to the online vehicle insurance system. Through system preprocessing, a standard and effective standard vehicle insurance document image is obtained. Then, based on the pre-trained OCR model and polygon fitting technology, text localization and segmentation are performed on the standard vehicle insurance document image to extract the text description therein. Through the keywords, themes, and context information of the text description, a preliminary classification is performed, and then it is fused with the visual features of the standard input image, the text features of the text description, and the result of the preliminary classification to obtain a vehicle insurance image fusion feature that more effectively represents the standard vehicle insurance document image. Finally, the vehicle insurance image fusion feature is matched with the preset vehicle insurance image template features, so as to effectively divide the standard vehicle insurance document image into the corresponding vehicle insurance document image category. To facilitate the subsequent system to input image information according to the input method corresponding to the vehicle insurance document image category.

[0050] In some alternative implementation manners of this embodiment, the steps of obtaining an input image and preprocessing the input image to obtain a standard input image include the following:

[0051] Perform gray-scale conversion and binarization on the input image to obtain a comparison input image;

[0052] In this embodiment, the input color image is converted into a gray-scale image. Each pixel in the gray-scale image has only one brightness value, usually ranging from 0 (black) to 255 (white), so as to make subsequent image processing more efficient. Binarization is to convert the gray-scale image into an image that only contains two pixel values (usually 0 and 255, that is, black and white). For example, the uploaded picture of the auto insurance document is converted into a black-and-white comparison input image through gray-scale conversion and binarization processing to improve the speed of subsequent image processing.

[0053] Perform denoising processing on the comparison input image to obtain a denoised input image;

[0054] In this embodiment, the median filter denoising algorithm is applied to the binarized image to remove noise points or small blocks in the image. In this embodiment, the noise may be caused by various interference factors during the user's image acquisition process. For example, when taking pictures of the accident scene during an auto insurance accident, due to signal interference at the scene, certain noise points appear in the obtained auto insurance accident image, and then the input image uploaded to the online auto insurance system has noise.

[0055] Perform skew detection on the denoised input image according to edge detection and Hough transform, and perform image correction according to the detection result to obtain the standard input image.

[0056] In this embodiment, edges in the image are identified through an edge detection algorithm (such as the Canny edge detector). This edge part is where the brightness changes significantly in the image, usually corresponding to the contour of an object. The Hough transform is used to detect straight lines or curves in the image. The Hough transform is a feature extraction technique suitable for detecting geometric shapes in images. In this embodiment, it is used to detect whether the image is skewed and the skewed angle. According to the results of edge detection and Hough transform, it is judged whether the image is skewed and the skewed angle is calculated. Then, correction methods such as image rotation or affine transformation are applied to adjust the skewed image to the horizontal or vertical direction to obtain the standard input image. For example, due to the shaking of the lens during the shooting of the uploaded auto insurance accident image, even if the lens is focused during the shooting process, it will still be skewed. This skew may be easily overlooked by the user when the current image is clear after shooting, resulting in the skewed situation of the input image.

[0057] In this embodiment, the input image is subjected to grayscale conversion and binarization to obtain a comparison input image; the comparison input image is denoised to obtain a denoised input image; tilt detection is performed on the denoised input image according to edge detection and Hough transform, and image correction is performed according to the detection result to obtain the standard input image. Thus, the input image can be effectively converted into a standard input image without noise and tilt, which is convenient for subsequent text localization and segmentation operations.

[0058] In some optional implementation manners of this embodiment, the pre-trained OCR model combines with the polygon fitting technology to perform text localization and segmentation on the standard input image, and the steps for extracting text information include the following:

[0059] Input the standard input image into the pre-trained OCR model to obtain an image probability map;

[0060] In this embodiment, the pre-processed standard input image is input into the trained OCR model. Among them, OCR can be regarded as a sequence recognition problem similar to speech recognition, and the end-to-end recognition of the entire line of text is realized based on the algorithm of the CRNN-based entire line recognition technology (CNN + LSTM + CTC) model. It has been trained with a large amount of text image data and can recognize the text in the image. After the OCR model processes the input image, it will output an image probability map. The image probability map is a two-dimensional array, where the value of each element represents the probability that the corresponding image position is a certain character. The probability map can contain multiple channels, and each channel corresponds to a different character or character category.

[0061] Perform connected component analysis on the image probability map to obtain image connected components;

[0062] In this embodiment, connected component analysis is performed on the image probability map to identify the pixel regions (i.e., connected components) that are connected to each other in the image. The connected components correspond to text blocks or characters in the image. For example, the uploaded vehicle insurance document image includes the information of the applicant and the insured (applicant / insured name, document type and number, contact information), vehicle information (license plate number, vehicle model, vehicle identification code, vehicle color), insurance information (insurance company name, insurance type, insurance amount, insurance period), other information (insurance premium, vehicle inspection and verification situation), etc. After the vehicle insurance document image is input into the pre-trained OCR model, the corresponding image probability map is obtained, and then the pixel regions with high probability in the image probability map are connected to obtain the image connected components where the text such as the information of the applicant and the insured, vehicle information, insurance information, and other information is located.

[0063] Perform polygon fitting on the image connected components to generate the minimum bounding polygon;

[0064] In this embodiment, a polygon fitting algorithm is adopted to perform polygon fitting on each identified connected region to generate a minimum bounding polygon, which can tightly enclose the connected region while minimizing unnecessary edges as much as possible. In this embodiment, the minimum circumscribed polygon is used as the polygon fitting algorithm. The algorithm steps of the minimum circumscribed polygon include: selecting the topmost, leftmost, or rightmost point in the contour point set as the starting point, where this starting point can be arbitrary. Starting from the starting point, traverse the remaining point set to find the next boundary point such that the area of the polygon formed by the newly added point and the existing point set is minimized (or the perimeter is minimized, depending on the algorithm design). This process needs to be iterated continuously until all points are included in the polygon. After finding the last boundary point, connect it to the starting point to form a closed polygon.

[0065] Verify and merge the minimum bounding polygon to obtain the text region;

[0066] In this embodiment, the generated minimum bounding polygon is verified to exclude those polygons that do not conform to the characteristics of the text region (such as being too small in area, irregular in shape, etc.). Then, adjacent polygons that may belong to the same text block are merged to form a larger text region.

[0067] Perform text segmentation based on the text region to obtain the text information.

[0068] In this embodiment, according to the boundaries determined by the text region, the text region is segmented into individual characters or words. The segmentation operation involves operations such as cropping and rotating the image to ensure that each character or word is correctly segmented.

[0069] In this embodiment, the standard input image is input into the pre-trained OCR model to obtain an image probability map; connected region analysis is performed on the image probability map to obtain image connected regions; polygon fitting is performed on the image connected regions to generate a minimum bounding polygon; the minimum bounding polygon is verified and merged to obtain the text region; text segmentation is performed based on the text region to obtain the text information. Thus, it effectively realizes the extraction of text information in the standard input image based on the pre-trained OCR model, facilitating the subsequent preliminary classification of the standard input image according to the extracted text information.

[0070] In some optional implementation manners of this embodiment, the steps of converting the text information into text information and performing a preliminary classification on the standard input image according to the keywords, themes, and context information of the text information to obtain a preliminary classification result include the following steps:

[0071] Organize the text information into structured text information according to regions;

[0072] In this embodiment, the extracted text information is organized according to its regions in the image to form structured text information. This step includes identifying the hierarchical relationships among paragraphs, sentences, and words, as well as their relative positions in the image, which can be achieved by a pre-trained convolutional neural network CNN.

[0073] Extract the keyword from the structured text information according to the keyword extraction method;

[0074] In this embodiment, a keyword extraction method is applied to extract important keywords from the structured text information. The keyword extraction method can use statistical features (such as word frequency, inverse document frequency), word embedding techniques (such as Word2Vec), or deep learning models (such as BERT) to evaluate the importance of words.

[0075] Perform semantic structure and semantic relationship analysis on the structured text information to obtain the theme and the context information;

[0076] In this embodiment, performing semantic structure and semantic relationship analysis on the text information includes identifying the logical relationships between sentences (such as causality, condition, parallelism, etc.), identifying the topic sentences and key information points, and constructing the semantic network of the text. Through semantic analysis, the theme of the text, that is, the core content discussed by the text, is determined. At the same time, the context information related to the theme is extracted, including details related to the theme, related events, etc. The above semantic structure and semantic relationship analysis can be achieved by a pre-trained BERT model.

[0077] Based on the keyword, the theme, and the context information, preliminarily classify the standard input image according to the preset classification rules to obtain the preliminary classification result.

[0078] In this embodiment, the preliminary classification result is output, that is, the category or label to which the standard input image belongs. The preset classification rules can be formulated based on a domain-specific knowledge base, predefined classification labels, or the prediction results of machine learning models (such as support vector machines, decision trees, neural networks, etc.). For example, extract keywords / themes from a car insurance document image: driving license, vehicle license, insurance policy; context information: the text content on the document is clear and contains the basic information of the vehicle and insurance information. The classification rule is to classify according to the type of document, such as driving license, vehicle license, insurance policy, etc. Then the preliminary classification result is that this car insurance document image belongs to the "insurance policy" category. Or, extract keywords / themes from a car insurance document image: collision, scratch, waterlogging; context information: the accident description part on the document clearly indicates that the vehicle has had a collision accident. The classification rule is to classify according to the nature of the accident, such as collision, scratch, waterlogging, fire, etc. Then the preliminary classification result is that this car insurance document image belongs to the "collision accident" category.

[0079] In this embodiment, the text information is organized into the structured text information according to regions; the keywords are extracted from the structured text information according to a keyword extraction method; the semantic structure and semantic relationship of the structured text information are analyzed to obtain the theme and the context information; and based on the keywords, the theme, and the context information, the standard input image is preliminarily classified according to a preset classification rule to obtain the preliminary classification result. Thus, it is effectively realized to accurately classify the standard input image according to the semantic information and semantic relationship contained in the text information, so as to provide an effective data basis for subsequent feature fusion processing.

[0080] In some optional implementation manners of this embodiment, the steps of extracting the visual features of the standard input image and the text features of the text information, and performing feature fusion on the visual features, the text features, and the preliminary classification result by using a hierarchical feature fusion method to obtain image fusion features include the following:

[0081] Feature extraction is performed on the standard input image according to a preset visual encoder to obtain the visual features;

[0082] In this embodiment, a preset visual encoder is used to perform feature extraction on the standard input image. This visual encoder can be a deep learning model, such as a convolutional neural network (CNN), which can extract discriminative visual features from the image.

[0083] Semantic feature extraction is performed on the text information to obtain the text features;

[0084] In this embodiment, semantic feature extraction is performed on the text information obtained from the previous processing. This includes using word embedding techniques (such as BERT, etc.) to convert the text into points in a high-dimensional vector space, so as to obtain the corresponding text features.

[0085] The preliminary classification result is converted into a result feature vector, and the visual features, the text features, and the result feature vector are fused to obtain a first fusion feature;

[0086] In this embodiment, the preliminary classification result (i.e., the category or label to which the standard input image belongs) is converted into a result feature vector. This can be achieved by mapping the classification label to a predefined vector space or by training a classifier to output a vector representing the classification result. For example, a predefined vector space is defined, which can be a high-dimensional sparse vector space, where each dimension corresponds to a possible classification label. The number of dimensions of the vector is equal to the number of all possible classification labels. For the preliminary classification result, a corresponding vector is created in the predefined vector space. In this vector, the dimension corresponding to the classification label is set to 1 (or a non-zero value), while all other dimensions are set to 0. Finally, after performing vector space mapping on the preliminary classification result, a sparse, high-dimensional feature vector is obtained, which uniquely represents the preliminary classification result, and this representation is the result feature vector. The fusion method for fusing visual features, text features, and the result feature vector can adopt weighted summation. By presetting weights for the visual features, text features, and the result feature vector, then multiplying each feature vector by its corresponding preset weight, and then adding the obtained weighted feature vectors, the final first fusion feature is obtained.

[0087] Semantic information is extracted from the first fusion feature based on a neural network to obtain a second fusion feature;

[0088] In this embodiment, a neural network (such as an attention mechanism network) is used to further process the first fusion feature to extract higher-level semantic information and obtain a second fusion feature. This second fusion feature contains richer and deeper semantic representations of the image and text information.

[0089] The first fusion feature and the second fusion feature are combined to obtain the image fusion feature.

[0090] In this embodiment, the first fusion feature and the second fusion feature are combined to obtain the final image fusion feature, where the combination method in the above feature combination step adopts feature splicing.

[0091] In this embodiment, the visual features are obtained by extracting features from the standard input image according to a preset visual encoder; the text features are obtained by extracting semantic features from the text information; the preliminary classification result is converted into a result feature vector, and the visual features, the text features, and the result feature vector are fused to obtain a first fused feature; semantic information is extracted from the first fused feature based on a neural network to obtain a second fused feature; and the first fused feature and the second fused feature are combined to obtain the image fusion feature. Thus, it is effectively realized to obtain an image fusion feature that more accurately expresses the features of the standard input image according to the visual features, text features, and result feature vector, so as to improve the accuracy of subsequent classification of the input image category.

[0092] In some optional implementation manners of this embodiment, the steps of calculating the similarity between the image fusion feature and a pre-stored image template feature to obtain a similarity result, and classifying the standard input image based on the similarity result to obtain the input image category include the following:

[0093] Calculate the cosine similarity between the image fusion feature and the image template feature to obtain the similarity result;

[0094] In this embodiment, the cosine similarity between the image fusion feature and the image template feature is calculated. Cosine similarity is a measure of the directional similarity between two vectors, and its value ranges from -1 to 1. When the directions of the two vectors are exactly the same, the cosine similarity is 1; when the directions are exactly opposite, the cosine similarity is -1; when the two vectors are orthogonal, the cosine similarity is 0.

[0095] Judge whether the similarity value corresponding to the similarity result is greater than or equal to a preset similarity threshold;

[0096] In this embodiment, the similarity threshold is a preset judgment value for determining whether two features are similar. The similarity threshold can be a decimal value or a percentage value, and the numerical type of the similarity value (i.e., decimal value or percentage value) corresponds to that of the similarity threshold. In this embodiment, the similarity threshold is preset to 80%, and it can be adjusted accordingly according to the actual situation.

[0097] If the similarity value corresponding to the similarity result is greater than or equal to the preset similarity threshold, then classify the standard input image into the input image category corresponding to the image template feature;

[0098] In this embodiment, if the calculated similarity value is greater than or equal to a preset similarity threshold, it is considered that the standard input image is similar enough to the image template feature. Therefore, the standard input image is classified into the input image category corresponding to the image template feature. The image template feature can represent a specific category of images and has a certain robustness to image variations within this category. For example, if the similarity value of the image template feature of the uploaded vehicle insurance document image for the compulsory traffic insurance document is greater than or equal to the corresponding similarity threshold, and the similarity value of the image template feature for the commercial insurance document is less than the corresponding similarity threshold, then the input vehicle insurance document image is classified as a compulsory traffic insurance document image to facilitate the system to enter data according to the corresponding entry template.

[0099] If the similarity value corresponding to the similarity result is less than the preset similarity threshold, the system will continue to calculate the similarity and perform threshold judgment between the image fusion feature and other image template features until all the image template features are traversed, or the similarity value between the image fusion feature and other image template features is greater than or equal to the similarity threshold.

[0100] In this embodiment, if the similarity value between the image fusion feature of the standard input image and a certain image template feature is less than the preset similarity threshold, the system will continue to calculate the similarity between the image fusion feature of the standard input image and the remaining other image template features. The system will compare each possible category one by one to find the most matching category. For each new image template feature, the system will calculate the similarity value and compare it with the preset similarity threshold again. This process will continue until the system has traversed all the image template features or found an image template feature with a similarity value greater than or equal to the threshold. If an image template feature with a similarity value greater than or equal to the threshold is found during the traversal, the system will classify the standard input image into the input image category corresponding to this template feature. If no template feature that meets the conditions is found after traversing all the image template features (i.e., all similarity values are less than the threshold), the system may classify the standard input image as an "unknown category" or "mismatched category", or take other processing measures according to the preset processing steps.

[0101] In this embodiment, the cosine similarity is calculated between the image fusion feature and the image template feature to obtain the similarity result; it is judged whether the similarity value corresponding to the similarity result is greater than or equal to the preset similarity threshold; if the similarity value corresponding to the similarity result is greater than or equal to the preset similarity threshold, the standard input image is classified into the input image category corresponding to the image template feature. Thus, it effectively realizes the accurate classification of the input image category to which the standard input image belongs, facilitating the subsequent entry or storage of the input image.

[0102] In some alternative implementation manners of this embodiment, before calculating the cosine similarity between the image fusion feature and the image template feature to obtain the similarity result, the following steps are further included:

[0103] Obtain a sample classification image;

[0104] In this embodiment, the sample classification image is a classified image corresponding to the input image. The sample classification image covers all classification categories. For example, the sample classification image covers all classification categories corresponding to vehicle insurance document images, and each category has a sufficient number of samples to ensure the accuracy of subsequent calculations.

[0105] Extract visual features and text features from the sample classification image respectively to obtain sample visual features and sample classification features;

[0106] In this embodiment, a deep learning model (such as a convolutional neural network CNN) can be used to extract key visual information in the image, such as edges, textures, shapes, and colors, when extracting visual features from each sample classification image. Natural language processing (NLP) techniques can be used to extract semantic information of the text, such as word embeddings, topic models, etc., when extracting features from the text of the sample classification image.

[0107] Fuse the sample visual features and the sample classification features to obtain sample fusion features;

[0108] In this embodiment, the extracted visual features and text features are fused to generate sample fusion features. The fusion method can adopt feature splicing.

[0109] Obtain classification category information, calculate the average value of the sample fusion features corresponding to each classification category information to obtain category feature codes;

[0110] In this embodiment, obtain the classification category information of each image from the sample data. This information usually exists in the form of labels and is used to indicate which category each sample classification image belongs to. For each classification category, calculate the average value of all its sample fusion features to generate the feature code of this category. This feature code represents the typical features of this category and can be used as a template for subsequent classification tasks.

[0111] Save the category feature codes as the image template features.

[0112] In this embodiment, save the calculated category feature codes as image template features for comparison with the features of the input image uploaded by the user to the system to achieve the classification purpose.

[0113] In this embodiment, sample classification images are obtained; visual feature extraction and text feature extraction are respectively performed on the sample classification images to obtain sample visual features and sample classification features; the sample visual features and the sample classification features are fused to obtain sample fusion features; classification category information is obtained, and the average value of the sample fusion features corresponding to each classification category information is calculated to obtain category feature codes; and the category feature codes are saved as the image template features. Thus, it is effectively realized to generate image template features that effectively represent the features of the sample classification images, so as to provide a reasonable and effective basis for the effective division of input images.

[0114] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through computer-readable instructions, and the computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0115] It should be understood that although the steps in the flowchart of the accompanying drawings are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps is not strictly limited in order, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0116] Further reference Figure 3 , as an implementation of the method shown above Figure 1 , the present application provides an embodiment of an input image category recognition device. This device embodiment corresponds to the method embodiment shown Figure 1 , and this device can be specifically applied to various electronic devices.

[0117] As Figure 3 shown, the input image category recognition device 600 described in this embodiment includes: an image processing module 601, a text extraction module 602, a preliminary classification module 603, a feature fusion module 604, and a category division module 605. Among them:

[0118] The image processing module 601 is configured to obtain an input image, preprocess the input image, and obtain a standard input image;

[0119] The text extraction module 602 is configured to perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with a polygon fitting technique, and extract text information;

[0120] The preliminary classification module 603 is configured to convert the text information into text data, and perform preliminary classification on the standard input image according to keywords, topics, and context information of the text data, to obtain a preliminary classification result;

[0121] The feature fusion module 604 is configured to extract visual features of the standard input image and text features of the text information, and perform feature fusion on the visual features, the text features, and the preliminary classification result by using a hierarchical feature fusion method, to obtain image fusion features;

[0122] The category division module 605 is configured to calculate a similarity between the image fusion features and pre-stored image template features, to obtain a similarity result, and perform category division on the standard input image based on the similarity result, to obtain an input image category.

[0123] By adopting the above input image category recognition device in this embodiment, it is possible to obtain an input image, preprocess the input image to obtain a standard input image; perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with a polygon fitting technique, and extract text information; convert the text information into text data, and perform preliminary classification on the standard input image according to keywords, topics, and context information of the text data, to obtain a preliminary classification result; extract visual features of the standard input image and text features of the text information, and perform feature fusion on the visual features, the text features, and the preliminary classification result by using a hierarchical feature fusion method, to obtain image fusion features; calculate a similarity between the image fusion features and pre-stored image template features, to obtain a similarity result, and perform category division on the standard input image based on the similarity result, to obtain an input image category. Thereby, accurate category recognition of the input image can be effectively achieved, and the reliability of the recognition result can be improved.

[0124] In some optional implementation manners of this embodiment, the image processing module 601 includes: a first processing unit, a second processing unit, and a third processing unit. Wherein:

[0125] The first processing unit is configured to perform grayscale conversion and binarization on the input image, to obtain a comparison input image;

[0126] The second processing unit is configured to perform denoising processing on the comparison input image to obtain a denoised input image;

[0127] The third processing unit is configured to perform skew detection on the denoised input image according to edge detection and Hough transform, and perform image correction according to the detection result to obtain the standard input image.

[0128] In this embodiment, by setting the image processing module 601 including the first processing unit, the second processing unit, and the third processing unit, the input image can be effectively converted into a standard input image without noise and skew, so as to facilitate subsequent text localization and segmentation operations.

[0129] In some alternative implementation manners of this embodiment, the text extraction module 602 includes: a model processing unit, an image connectivity unit, a graphic fitting unit, a region arrangement unit, and a text segmentation unit. Among them:

[0130] The model processing unit is configured to input the standard input image into the pre-trained OCR model to obtain an image probability map;

[0131] The image connectivity unit is configured to perform connected component analysis on the image probability map to obtain an image connected component;

[0132] The graphic fitting unit is configured to perform polygon fitting on the image connected component to generate a minimum bounding polygon;

[0133] The region arrangement unit is configured to verify and merge the minimum bounding polygon to obtain a text region;

[0134] The text segmentation unit is configured to perform text segmentation based on the text region to obtain the text information.

[0135] In this embodiment, by setting the text extraction module 602 including the model processing unit, the image connectivity unit, the graphic fitting unit, the region arrangement unit, and the text segmentation unit, the text information in the standard input image can be effectively extracted based on the pre-trained OCR model, so as to facilitate subsequent preliminary classification of the standard input image according to the extracted text information.

[0136] In some alternative implementation manners of this embodiment, the preliminary classification module 603 includes: a text organization unit, a keyword extraction unit, a context analysis unit, and a preliminary classification unit. Among them:

[0137] The text organization unit is configured to organize the text information into the structured text information according to regions;

[0138] The keyword extraction unit is configured to extract the keywords from the structured text information according to a keyword extraction method;

[0139] The context analysis unit is configured to perform semantic structure and semantic relationship analysis on the structured text information to obtain the theme and the context information;

[0140] The preliminary classification unit is configured to perform preliminary classification on the standard input image based on the keywords, the theme, and the context information according to a preset classification rule to obtain the preliminary classification result.

[0141] In this embodiment, by setting a preliminary classification module 603 including a text organization unit, a keyword extraction unit, a context analysis unit, and a preliminary classification unit, it is effectively realized to accurately classify the standard input image according to the semantic information and semantic relationship contained in the text information, so as to provide an effective data basis for subsequent feature fusion processing.

[0142] In some optional implementation manners of this embodiment, the feature fusion module 604 includes: a visual feature extraction unit, a text feature extraction unit, a first feature fusion unit, a secondary feature extraction unit, and a second feature fusion unit. Among them:

[0143] The visual feature extraction unit is configured to extract features from the standard input image according to a preset visual encoder to obtain the visual features;

[0144] The text feature extraction unit is configured to perform semantic feature extraction on the text information to obtain the text features;

[0145] The first feature fusion unit is configured to convert the preliminary classification result into a result feature vector, and fuse the visual features, the text features, and the result feature vector to obtain a first fusion feature;

[0146] The secondary feature extraction unit is configured to perform semantic information extraction on the first fusion feature based on a neural network to obtain a second fusion feature;

[0147] The second feature fusion unit is configured to combine the first fusion feature and the second fusion feature to obtain the image fusion feature.

[0148] In this embodiment, by setting a feature fusion module 604 including a visual feature extraction unit, a text feature extraction unit, a first feature fusion unit, a secondary feature extraction unit, and a second feature fusion unit, it is effectively realized to obtain an image fusion feature that more accurately expresses the features of the standard input image according to the visual features, text features, and result feature vector, so as to improve the accuracy of subsequent input image category division.

[0149] In some alternative implementation manners of this embodiment, the category division module 605 includes: a similarity calculation unit, a similarity judgment unit, a first processing unit, and a second processing unit.

[0150] Among them:

[0151] The similarity calculation unit is configured to calculate the cosine similarity between the image fusion feature and the image template feature to obtain the similarity result;

[0152] The similarity judgment unit is configured to judge whether the similarity value corresponding to the similarity result is greater than or equal to a preset similarity threshold;

[0153] The first processing unit is configured to, if the similarity value corresponding to the similarity result is greater than or equal to the preset similarity threshold, divide the standard input image into an input image category corresponding to the image template feature;

[0154] The second processing unit is configured to, if the similarity value corresponding to the similarity result is less than the preset similarity threshold, continue to calculate the similarity and perform threshold judgment on the image fusion feature and other image template features until all the image template features are traversed, or the similarity value between the image fusion feature and other image template features is greater than or equal to the similarity threshold.

[0155] In this embodiment, by setting the category division module 605 including a similarity calculation unit, a similarity judgment unit, a first processing unit, and a second processing unit, the accurate division of the input image category to which the standard input image belongs is effectively realized, so as to facilitate the subsequent input or storage of the input image.

[0156] In some alternative implementation manners of this embodiment, before the category division module 605, there are further included: a sample image acquisition unit, a sample feature extraction unit, a sample feature fusion unit, a feature code generation unit, and a feature storage unit. Among them:

[0157] The sample image acquisition unit is configured to acquire a sample classification image;

[0158] The sample feature extraction unit is configured to respectively perform visual feature extraction and text feature extraction on the sample classification image to obtain a sample visual feature and a sample classification feature;

[0159] The sample feature fusion unit is configured to perform feature fusion on the sample visual feature and the sample classification feature to obtain a sample fusion feature;

[0160] The feature code generation unit is configured to obtain classification category information, calculate the average value of the sample fusion features corresponding to each piece of classification category information, and obtain a category feature code;

[0161] The feature storage unit is configured to store the category feature code as the image template feature.

[0162] In this embodiment, by setting a sample image acquisition unit, a sample feature extraction unit, a sample feature fusion unit, a feature code generation unit, and a feature storage unit before the category division module 605, it is effectively realized to generate an image template feature that effectively represents its features according to the sample classification image, so as to provide a reasonable and effective basis for the effective division of the input image.

[0163] To solve the above technical problems, an embodiment of the present application also provides a computer device. For details, please refer to Figure 4 , Figure 4 which is the basic structural block diagram of the computer device in this embodiment.

[0164] The computer device 7 includes a memory 71, a processor 72, and a network interface 73 that are communicatively connected to each other through a system bus. It should be noted that only the computer device 7 with components 71-73 is shown in the figure, but it should be understood that it is not required to implement all the shown components, and more or fewer components can be alternatively implemented. Among them, those skilled in the art of the present technology can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0165] The computer device can be a desktop computer, a notebook, a palm computer, a cloud server, and other computing devices. The computer device can perform human-computer interaction with the user through a keyboard, a mouse, a remote control, a touchpad, a voice control device, or other means.

[0166] The memory 71 includes at least one type of readable storage medium, which includes flash memory, hard disk, multimedia card, card-type memory (such as SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 71 may be an internal storage unit of the computer device 7, such as the hard disk or memory of the computer device 7. In other embodiments, the memory 71 may also be an external storage device of the computer device 7, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, FlashCard, etc. equipped on the computer device 7. Of course, the memory 71 may also include both the internal storage unit and the external storage device of the computer device 7. In this embodiment, the memory 71 is generally used to store the operating system and various application software installed in the computer device 7, such as computer-readable instructions for the input image category recognition method. In addition, the memory 71 may also be used to temporarily store various data that have been output or will be output.

[0167] In some embodiments, the processor 72 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 72 is generally used to control the overall operation of the computer device 7. In this embodiment, the processor 72 is used to run the computer-readable instructions stored in the memory 71 or process data, such as running the computer-readable instructions for the input image category recognition method.

[0168] The network interface 73 may include a wireless network interface or a wired network interface, and the network interface 73 is generally used to establish a communication connection between the computer device 7 and other electronic devices.

[0169] In this embodiment, by using the above computer device, it is possible to obtain an input image, preprocess the input image to obtain a standard input image; perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with polygon fitting technology to extract text information; convert the text information into text data, and perform a preliminary classification on the standard input image according to the keywords, themes, and context information of the text data to obtain a preliminary classification result; extract the visual features of the standard input image and the text features of the text data, and use a hierarchical feature fusion method to fuse the visual features, the text features, and the preliminary classification result to obtain image fusion features; calculate the similarity between the image fusion features and pre-stored image template features to obtain a similarity result, and perform category division on the standard input image based on the similarity result to obtain the input image category. Thus, accurate category recognition of the input image can be effectively achieved, and the reliability of the recognition result can be improved.

[0170] The present application also provides another implementation manner, that is, to provide a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor, so that the at least one processor executes the steps of the input image category recognition method as described above.

[0171] In this embodiment, by using the above computer-readable storage medium, it is possible to obtain an input image, preprocess the input image to obtain a standard input image; perform text localization and segmentation on the standard input image based on a pre-trained OCR model combined with polygon fitting technology to extract text information; convert the text information into text data, and perform a preliminary classification on the standard input image according to the keywords, themes, and context information of the text data to obtain a preliminary classification result; extract the visual features of the standard input image and the text features of the text data, and use a hierarchical feature fusion method to fuse the visual features, the text features, and the preliminary classification result to obtain image fusion features; calculate the similarity between the image fusion features and pre-stored image template features to obtain a similarity result, and perform category division on the standard input image based on the similarity result to obtain the input image category. Thus, accurate category recognition of the input image can be effectively achieved, and the reliability of the recognition result can be improved.

[0172] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0173] Obviously, the above-described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The accompanying drawings show preferred embodiments of the present application, but do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosed content of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing specific embodiments or equivalently replace some of the technical features. Any equivalent structure directly or indirectly using the content of the specification and drawings of the present application in other related technical fields is similarly within the scope of the patent protection of the present application.

[0174] The non-company software tools or components appearing in the embodiments of the present application are only for illustrative introduction and do not represent actual use.

Claims

1. A method for identifying the category of an input image, characterized in that: The steps include: Acquire an input image, and preprocess the input image to obtain a standard input image; Based on the pre-trained OCR model combined with polygon fitting technology, the text is located and segmented on the standard input image to extract text information; Converting the text information into text information, and preliminarily classifying the standard input image according to the keywords, themes, and context information of the text information to obtain a preliminary classification result; Extracting visual features of the standard input image and text features of the text information, and fusing the visual features, the text features, and the preliminary classification results using a hierarchical feature fusion method to obtain image fusion features; The image fusion feature and the pre-stored image template feature are subjected to similarity calculation to obtain a similarity result, and the standard input image is classified into categories based on the similarity result to obtain an input image category.

2. The input image category recognition method according to claim 1, characterized in that: The step of obtaining an input image and preprocessing the input image to obtain a standard input image specifically includes: Performing grayscale conversion and binarization on the input image to obtain a contrast input image; Performing denoising processing on the contrast input image to obtain a denoised input image; The denoised input image is subjected to tilt detection according to edge detection and Hough transform, and image correction is performed according to the detection result to obtain the standard input image.

3. The input image category recognition method according to claim 1, characterized in that: The steps of locating and segmenting text on the standard input image based on the pre-trained OCR model combined with polygon fitting technology to extract text information specifically include: Inputting the standard input image into the pre-trained OCR model to obtain an image probability map; Performing connected domain analysis on the image probability map to obtain an image connected domain; Performing polygon fitting on the connected domain of the image to generate a minimum enclosing polygon; Verifying and merging the minimum enclosing polygons to obtain a text area; Perform text segmentation based on the text area to obtain the text information.

4. The input image category recognition method according to claim 1, characterized in that: The step of converting the text information into text information, and preliminarily classifying the standard input image according to the keywords, themes, and context information of the text information to obtain a preliminary classification result specifically includes: Organizing the text information into structured text information according to regions; Extracting the keywords from the structured text information according to a keyword extraction method; Performing semantic structure and semantic relationship analysis on the structured text information to obtain the subject and the context information; Based on the keywords, the subject, and the context information, the standard input image is preliminarily classified according to preset classification rules to obtain the preliminary classification result.

5. The input image category recognition method according to claim 1, characterized in that: The step of extracting the visual features of the standard input image and the text features of the text information, and fusing the visual features, the text features, and the preliminary classification results using a hierarchical feature fusion method to obtain image fusion features specifically includes: Extracting features from the standard input image according to a preset visual encoder to obtain the visual features; Extracting semantic features from the text information to obtain the text features; Converting the preliminary classification result into a result feature vector, and fusing the visual feature, the text feature, and the result feature vector to obtain a first fused feature; Extracting semantic information from the first fusion feature based on a neural network to obtain a second fusion feature; The first fusion feature and the second fusion feature are combined to obtain the image fusion feature.

6. The input image category recognition method according to claim 1, characterized in that: The step of calculating the similarity between the image fusion feature and the pre-stored image template feature to obtain a similarity result, and classifying the standard input image based on the similarity result to obtain an input image category specifically includes: Performing cosine similarity calculation on the image fusion feature and the image template feature to obtain the similarity result; Determine whether the similarity value corresponding to the similarity result is greater than or equal to a preset similarity threshold; If the similarity value corresponding to the similarity result is greater than or equal to a preset similarity threshold, the standard input image is divided into input image categories corresponding to the image template features.

7. The input image category recognition method according to claim 1, characterized in that: Before the step of calculating the cosine similarity of the image fusion feature and the image template feature to obtain the similarity result, the following steps are also included: Get sample classification images; Performing visual feature extraction and text feature extraction on the sample classification image respectively to obtain sample visual features and sample classification features; Performing feature fusion on the sample visual feature and the sample classification feature to obtain a sample fusion feature; Obtaining classification category information, calculating the average value of sample fusion features corresponding to each of the classification category information, and obtaining a category feature code; The category feature code is saved as the image template feature.

8. An input image category recognition device, characterized in that: include: An image processing module, used for acquiring an input image, and preprocessing the input image to obtain a standard input image; A text extraction module, used to locate and segment text in the standard input image based on a pre-trained OCR model combined with polygon fitting technology, and extract text information; A preliminary classification module, used to convert the text information into text information, and to perform preliminary classification on the standard input image according to the keywords, themes, and context information of the text information to obtain a preliminary classification result; A feature fusion module, used to extract the visual features of the standard input image and the text features of the text information, and to fuse the visual features, the text features and the preliminary classification results using a hierarchical feature fusion method to obtain an image fusion feature; The category classification module is used to calculate the similarity between the image fusion feature and the pre-stored image template feature to obtain a similarity result, and to classify the standard input image based on the similarity result to obtain an input image category.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and the processor implements the steps of the input image category recognition method according to any one of claims 1 to 7 when executing the computer-readable instructions.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the steps of the input image category recognition method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Asset statistical method and device, medium and equipment

    CN120893424A

  • Heterogeneous insurance policy image information extraction method, system, equipment and medium

    CN121640500A