Asset information intelligent completion method and system fusing multi-modal large model
By integrating a multimodal large model's visual-text-prior knowledge architecture, improving U-Net and bidirectional LSTM networks, and combining Transformer for intelligent asset label completion, the problem of low recognition accuracy of traditional asset management software when labels are worn or soiled in complex environments is solved, enabling real-time and accurate query of equipment information.
Patent Information
- Application Number
- CN202511136108.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-14
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-08-14
AI Technical Summary
Traditional asset management software suffers from low accuracy in recognizing worn or soiled tags in complex industrial environments, lacks intelligent reasoning and completion capabilities, and suffers from insufficient fusion of multi-source data, resulting in inadequate retrieval reliability and failing to meet enterprises' needs for real-time and accurate access to equipment information.
We employ a multimodal large model that integrates visual, textual, and prior knowledge architecture. We use an improved U-Net network for image segmentation and feature extraction, a bidirectional LSTM network for character sequence analysis, and Transformer for contextual reasoning. We dynamically adjust the multimodal feature fusion weights and activate a database fuzzy matching mechanism to ensure recognition accuracy.
It can stably process asset tag recognition in complex scenarios, improve the accuracy and reliability of tag recognition, and automatically infer and complete the tag when it is damaged, ensuring real-time and accurate query of equipment information.
Smart Images

Figure CN120747981B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of artificial intelligence, and particularly relates to an asset information intelligent completion method and system fusing a multi-modal large model. BACKGROUND
[0002] The statements in this section merely provide background information related to the present application and do not necessarily constitute the prior art.
[0003] With the development of digital asset management, asset management software has become a key tool for enterprises to efficiently manage assets. Due to the complex background of the industrial environment, real factors such as shooting angle and light can cause retrieval failure. At the same time, the traditional asset retrieval excessively relies on the completeness of the label. When the label is physically worn, stained, covered or partially missing, the accuracy of recognition will drop sharply. In addition, the existing scheme can only mechanically recognize visible characters and cannot intelligently infer and complete according to the context information such as asset type and location distribution. Moreover, it excessively relies on single coding information and lacks multi-source data fusion, and does not effectively combine the structured information of the existing asset database of the enterprise, resulting in insufficient reliability of the final retrieval and difficulty in meeting the needs of enterprises for real-time and accurate calling of device information. SUMMARY
[0004] In order to solve at least one technical problem in the background art, the present application provides an asset information intelligent completion method and system fusing a multi-modal large model, which fuses a multi-modal architecture of vision-text-prior knowledge and can stably process asset label recognition tasks in various complex scenarios.
[0005] In order to achieve the above-mentioned purpose, the present application adopts the following technical solutions:
[0006] The first aspect of the present application provides an asset information intelligent completion method fusing a multi-modal large model, comprising the following steps:
[0007] The obtained asset nameplate image is preprocessed and image enhanced to obtain an enhanced asset nameplate image;
[0008] The enhanced asset nameplate image is segmented to obtain a segmented asset nameplate image;
[0009] Multi-modal feature extraction is performed on the segmented asset nameplate image to obtain multi-modal features, the fusion weight of the multi-modal features is dynamically adjusted according to the missing degree of the image, and the fused multi-modal features are obtained by combining the fusion weight of the multi-modal features and the multi-modal features;
[0010] Based on the fused multi-modal features and the completion reasoning model, a complete image and text content are obtained.
[0011] Further, the obtained asset tag image is preprocessed and image enhanced, including:
[0012] The asset tag image is subjected to grayscale processing, the pixel values of the grayscale image are scaled, and the size of the input image is adjusted;
[0013] The preprocessed image is subjected to image enhancement, similar oil stains and textures are added to the training image in terms of data enhancement; the image is randomly occluded in a label missing manner, different occlusion area ratios and occlusion shapes are set, and finally the enhanced asset tag image is obtained.
[0014] Further, the improved U-Net network is used to segment the enhanced asset tag image, including: inputting the enhanced asset tag image into an encoder, using convolutional layers and pooling layers to extract image features, adding an attention mechanism SE before the jump connection at the end of the U-Net encoder, calibrating the feature map output by the encoder and then inputting it into the decoder, adding a residual block after each upsampling layer of the decoder to reconstruct the features, and setting an edge detection branch at the last layer of the decoder, using the high-dimensional feature map at the last layer of the U-Net decoder as an input variability convolution kernel, learning the coordinate offset in the high-dimensional feature map using parallel convolution, adding one or more residual blocks after the variability convolution to generate an edge map representing the probability of each pixel in the image belonging to the edge, and obtaining the final segmentation result.
[0015] Further, the segmented asset tag image is subjected to multi-modal extraction to obtain multi-modal features, including:
[0016] Based on the segmented asset tag image and the visual branch, local texture features are extracted;
[0017] The sequence features of the recognized characters are analyzed using a bidirectional LSTM network, including: obtaining the character sequence using OCR recognition, mapping each character to a word embedding vector, inputting the word embedding vector into the bidirectional LSTM network for forward LSTM and reverse LSTM, merging the forward LSTM and reverse LSTM outputs to obtain merged features, and fusing the merged features and prior knowledge embedding to obtain the final text feature vector representation.
[0018] Further, the fusion weight of the multi-modal features is determined by the image missing degree, and the calculation formula is:
[0019] ,
[0020] wherein, To score the image missing degree, the missing degree of the image is obtained according to the damage rate of the segmentation mask, which is obtained by the number of missing pixels / total number of pixels, wherein 0 represents complete, 1 represents severe missing, w is a weight parameter, and b is a bias term.
[0021] Further, based on the fused multi-modal features and the completion reasoning model, a complete image and text content are obtained, including:
[0022] Based on the fused multi-modal features, a primary completion result is obtained.
[0023] Based on the primary completion of the character image, a semantic middle-level completion is performed to obtain a semantically completed character image.
[0024] It is determined whether the confidence of the recognition result of the semantically completed character image obtained through the middle-level completion is higher than a set threshold value, and if it is lower than the set confidence threshold value, a database retrieval mechanism is activated, the recognition result is fuzzy matched with an asset database, a plurality of candidate results are retrieved and sorted, and a final completion result is selected according to the sorted result.
[0025] Further, the primary completion based on the fused multi-modal features includes:
[0026] In combination with the multi-modal features, based on character morphology analysis, the missing part is inferred through stroke continuity of adjacent characters to obtain a morphology repair result.
[0027] A non-local mean filtering algorithm is used to output a texture repair result through input of a segmented image mask and an original image.
[0028] The morphology repair result and the texture repair result are fused to obtain a primary completion result.
[0029] Further, the semantic middle-level completion based on the primary completion of the character image includes:
[0030] The character sequence recognized based on the primary completion of the character image and the corresponding context information thereof.
[0031] The character sequence and the corresponding context information thereof are converted into features suitable for a Transform processing format, the relationship between characters is extracted based on a Transform encoder, business rule constraints are added, the input sequence is deeply processed by the Transform encoder, the complex relationship between characters is captured by a multi-head self-attention mechanism, and the most likely complete character sequence is generated as a final completion suggestion.
[0032] The second aspect of the present application provides an asset information intelligent completion system integrating a multi-modal large model, including:
[0033] an image enhancement module, configured to pre-process and image enhance the acquired asset nameplate image to obtain an enhanced asset nameplate image;
[0034] an image segmentation module, configured to segment the enhanced asset nameplate image to obtain a segmented asset nameplate image;
[0035] a feature extraction module, configured to extract multi-modal features from the segmented asset nameplate image, dynamically adjust a fusion weight of the multi-modal features according to a missing degree of the image, and obtain fused multi-modal features by combining the fusion weight of the multi-modal features and the multi-modal features;
[0036] an information completion module, configured to obtain a complete image and text content based on the fused multi-modal features and a completion inference model.
[0037] A third aspect of the present application provides a computer device.
[0038] A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the asset information intelligent completion method of the fused multi-modal large model when executing the program.
[0039] Compared with the prior art, the present application has the following advantages:
[0040] 1. The present application fully acquires image data under various complex conditions, fuses a multi-modal architecture of vision-text-prior knowledge, and performs completion inference on missing information based on the multi-modal architecture to obtain complete information, which can stably process asset label recognition tasks under various complex scenes.
[0041] 2. The present application improves the image segmentation network structure by adding an attention mechanism SE (Squeeze-and-Excitation) and an edge detection branch; the attention mechanism SE is used to enhance the feature expression of key regions, residual connections are used to improve the network depth and stability, and the network can recognize the contour features of various damaged labels after training, and can accurately segment even if there is a large area of surface damage.
[0042] 3. The present application proposes a three-level completion inference model, if the primary or intermediate completion has accurately recognized complete information, the high-level completion is skipped, and if the recognition result confidence is low, the database fuzzy matching mechanism is activated to ensure the reliability of the system output.
[0043] Advantages of the present application's additional aspects will be readily apparent to those skilled in the art from the following description, merely by way of illustration of BRIEF DESCRIPTION OF DRAWINGS
[0044] The accompanying drawings, which form a part of this specification, are included to provide a further understanding of the application, and are incorporated by reference herein. The drawings are not necessarily to scale, the embodiments of the application, and the descriptions thereof, are not intended to limit the scope of the application, and are included merely for the purposes of explanation.
[0045] Figure 1 is a process flow diagram of the asset information intelligent completion method of the fusion multi-modal large model provided by the embodiment of the application;
[0046] Figure 2 is a schematic diagram of the image segmentation process provided by the embodiment of the application;
[0047] Figure 3 is a block diagram of the asset information intelligent completion system of the fusion multi-modal large model provided by the embodiment of the application;
[0048] Figure 4 is a schematic diagram of the computer device structure provided by the embodiment of the application. DETAILED DESCRIPTION
[0049] The embodiments of the present application are described below with reference to the accompanying drawings. In the following description, reference is made to the accompanying drawings that form a part hereof, and in which are shown by way of illustration specific aspects in which embodiments of the application can be practiced. It is to be understood that other aspects can be utilized and that structural or logical changes can be made without departing from the scope of the present application. For instance, it is to be understood that the disclosure made in connection with a described method can also hold true for a corresponding device or system configured to perform the method and vice versa. For example, if one or more specific method steps are described, a corresponding device can include one or more units, e.g., functional units, to perform the described one or more method steps, even though the one or more units are not explicitly described or illustrated in the figures. On the other hand, if a specific apparatus is described, a corresponding method can include one step to perform the functionality of the one or more units, even though the one or more steps are not explicitly described or illustrated in the figures. Further, it is understood that features described with respect to one example embodiment or aspect can be combined with features described with respect to another example embodiment or aspect.
[0050] As mentioned in the background of the present application, with the development of digital asset management, asset management software has become a key tool for enterprises to efficiently manage assets. Due to the complex background of industrial environment, real factors such as shooting angle and light will cause retrieval failure, at the same time, the traditional asset retrieval excessively depends on the integrity of the label image, when the label image appears physical wear, stain covering or partial loss, the accuracy of recognition will drop sharply, in addition, the existing scheme can only mechanically identify visible characters, cannot intelligently infer and complete according to the context information such as asset type and position distribution, and excessively relies on single coding information, multi-source data fusion is insufficient, and the structured information of the existing asset database of the enterprise is not effectively combined, resulting in insufficient reliability of the final retrieval, which is difficult to meet the demand of enterprises for real-time and accurate calling of equipment information.
[0051] The present application is aimed at the application scene of asset management software, fully combines the powerful semantic understanding and feature learning ability of multi-modal large model, realizes accurate extraction and completion of text information of equipment pictures, effectively eliminates the similarity interference of worn-out text, and efficiently queries equipment information without the uniqueness of asset coding.
[0052] Specifically includes the following aspects:
[0053] 1. Based on the improved visual processing flow, the system can effectively handle common imaging problems such as tilt, reflection, shadow, etc., and maintain stable recognition performance under various lighting conditions. It is particularly optimized for industrial environment to ensure accurate extraction of label information in complex background.
[0054] 2. Integrate enterprise asset database information, automatically associate related business data during identification process. When encountering incomplete information, it can make reasonable inference based on device type, storage location, etc. Context, significantly improve the processing effect of incomplete labels.
[0055] 3. Adopting modular design architecture, supporting continuous optimization of recognition model through adding sample data. When enterprises introduce new label formats, they can quickly adapt through supplementary training to maintain the long-term applicability of the system.
[0056] Figure 1 The asset information intelligent completion method fusing multi-modal large model provided by the embodiment of the present application includes the following steps:
[0057] S101: Preprocess and image enhance the obtained asset nameplate image to obtain an enhanced asset nameplate image;
[0058] Before segmenting the asset tag image, the asset tag image is first subjected to grayscale processing, which can reduce the influence of lighting, and then the pixel values of the grayscale image are scaled to the range of [0, 1] to adapt to the input requirements of the neural network. After adjusting the size of the input image, the preprocessing operation is completed.
[0059] Image enhancement is performed on the preprocessed image, and similar oil stains and textures are added to the training image in terms of data enhancement; and a part of the image is randomly occluded to simulate the case of missing labels.
[0060] Specifically, when adding similar oil stains and textures to the training image, existing algorithms can be used, for example, Perlin noise is used to simulate the fluid diffusion pattern of oil stains, and then the texture is extracted from a real oil stain picture and fused into the label image by Poisson Blending.
[0061] When simulating the case of missing labels, random deletion of local line segments can be used to simulate stroke breakage, and after a random polygon is used to occlude an area, morphological erosion is performed on the occlusion boundary.
[0062] In order to enhance the diversity of samples, different occlusion area ratios and occlusion shapes can be set, and finally an asset tag image training data set is obtained.
[0063] S102: Segmenting the asset tag image to obtain a segmented asset tag image.
[0064] Figure 2 The image segmentation network structure provided by the embodiments of the present application uses an improved U-Net network. The traditional U-Net network includes an encoder, a decoder, and a skip connection. The encoder is composed of a series of convolutional layers, activation functions (such as ReLU), and max-pooling layers. The decoder is composed of a series of up-sampling, convolutional layers, and activation functions (such as ReLU). The low-level features in the encoder are combined with the high-level features in the decoder through the skip connection, so that the network can capture multi-scale information.
[0065] The U-Net network is not good at segmenting small targets or label edges, and is sensitive to noise, pollution, and changes in lighting. Considering that the text, numbers, and symbols on the asset tag are usually small targets, with fine strokes and sharp edges, poor segmentation will lead to OCR recognition failure. Moreover, the asset tag itself may have a complex shape (such as a rounded corner, a mounting hole, or a specific border), and its accurate contour is important for positioning and subsequent processing. In addition, physical pollution may exist in part of the asset tag, and environmental noise, imaging noise, or on-site lighting conditions are difficult to control, which may cause overexposure or underexposure of part of the asset tag.
[0066] In this embodiment, the attention mechanism SE (Squeeze-and-Excitation) and the edge detection branch are introduced; the encoder extracts image features using convolutional layers and pooling layers, at the end of the encoder of the U-Net, the attention mechanism SE is introduced before the jump connection, the feature map output by the encoder is calibrated and then input into the decoder, a residual block is added after each upsampling layer of the decoder to reconstruct the features and improve the focus of the model on key characteristics to identify the presence or absence of the missing part, and at the same time, an edge detection branch is set at the output end of the decoder to improve the irregular shape of the damaged edge.
[0067] The attention mechanism SE is introduced to enhance the feature expression of the key region, which includes the following steps:
[0068] First, the shape, structure and texture features of the asset nameplate image are extracted respectively to perform global average pooling to generate a channel descriptor to capture the global distribution of each channel.
[0069] Then, a fully connected network is used to learn the nonlinear dependence between channels to output the importance weight of each channel, and in this embodiment, the importance weight of each channel is a scaling coefficient between 0 and 1; the coefficient close to 1 indicates that the channel information is crucial in the current image area and should be amplified; the coefficient close to 0 indicates that the channel information is redundant or has strong interference and should be suppressed.
[0070] The learned importance weight of each channel is multiplied by the corresponding channel of the original feature map to obtain a weighted feature map.
[0071] Through the above steps, the SE module can enhance the expression of key information channels such as small text strokes, sharp edges of nameplates, and small icons by automatically learning the importance weight of the feature channel, and suppress the interference of background or unimportant channels. At the same time, the feature channels contaminated by oil stains, rust, strong light or shadows are suppressed, and the channels that are relatively stable under light changes and can better represent the essence of the nameplate are enhanced, improving the robustness of the network to harsh imaging conditions.
[0072] The attention mechanism SE introduced in the present application enhances the feature expression of the key region, uses residual connections to improve the depth and stability of the network, and through training of the network, the contour features of various damaged nameplates can be recognized, and even if there is a large area of contamination on the surface of the nameplate, the nameplate can be accurately segmented.
[0073] In this embodiment, the edge detection branch includes variable convolution and residual blocks. The variable convolution shares the high-dimensional features output by the backbone network and learns the coordinate offset through parallel convolution. After the variable convolution, one or more residual blocks are added to obtain the final segmentation result.
[0074] Specifically, the high-dimensional feature map of the last layer of the U-Net decoder is used as input, integrating global and local information to determine the location and approximate shape of the target in the image. Parallel convolutional layers are used to automatically infer the spatial coordinate offset of each convolutional kernel sampling point at each spatial location from the input high-dimensional feature map. The learned spatial coordinate offset Superimposed onto the fixed sampling coordinates of a standard convolution kernel, a variable convolution kernel is formed. The sampling point position of the convolution kernel can be dynamically adjusted, no longer limited to a regular rectangular grid. The convolution kernel "deforms" to adapt to the actual geometry of the target. The advantage of this convolution method is that it allows the model to dynamically adjust its receptive field to adapt to target edges of different scales and shapes, especially irregular edges caused by dirt, damage, etc.
[0075] After the deformable convolution, one or more residual blocks are added to further refine and deepen the edge feature representation that has been initially adjusted by the deformable convolution. The residual blocks help alleviate the gradient vanishing problem in deep networks and allow the model to learn the residual mapping between the input and output more effectively, which is especially important for the extraction of fine edges. After the residual blocks, the edge detection branch produces an edge mapping that represents the probability that each pixel in the image belongs to an edge, as the output.
[0076] In summary, introducing variable convolutions through the edge detection branch can adaptively adjust its receptive field to better capture challenging edge features, such as damaged or occluded parts, thereby enhancing the ability to capture complex edges. The application of residual blocks not only helps solve the difficulties in training deep networks but also improves the model's ability to process detailed information, making it more stable in the face of noise, dirt, and changes in lighting. Since variable convolutions allow the convolution kernel to dynamically adjust its sampling position according to the input feature map, the edge detection branch can more accurately locate edge positions in complex backgrounds, thereby improving the overall edge segmentation accuracy.
[0077] S103: Perform multimodal extraction on the segmented asset nameplate image to obtain multimodal features. Dynamically adjust the fusion weight of the multimodal features according to the degree of image loss. Combine the fusion weight of the multimodal features and the multimodal features to obtain the fused multimodal features.
[0078] The feature extraction network in the embodiment includes a visual branch and a text branch. The visual branch adopts a lightweight ViT model to extract local texture features, and particularly enhances the perception ability of damaged edges. The text branch uses a bidirectional LSTM network to analyze the sequence features of the recognized characters, and simultaneously accesses an enterprise asset database to generate prior knowledge embedding. The two features are adaptively weighted in a fusion module, and the weight is dynamically adjusted according to the label damage degree, so as to ensure that the optimal feature combination can be maintained under different damage conditions.
[0079] The method comprises the following steps:
[0080] S301, extracting local texture features and damaged edge perception features based on the segmented asset nameplate image and the visual branch ;
[0081] In the embodiment, a lightweight ViT model can be used to extract local texture features, including stroke details and rust texture features. The Patch Embedding and Self-Attention mechanism of ViT can focus on local regions, are better at modeling long-range dependencies than CNN, and are suitable for capturing the relevance of discrete damaged regions.
[0082] Meanwhile, edge prior is introduced into the position encoding or attention mask of ViT. The edge probability map in the segmentation stage is used to guide the model to pay attention to the boundary region.
[0083] S302, using a bidirectional LSTM network to analyze the sequence features of the recognized characters, specifically comprising:
[0084] S3021, obtaining a character sequence by using OCR recognition ;
[0085] S3022, mapping each character to a word embedding vector ;
[0086] S3023, inputting the word embedding vector into a bidirectional LSTM network for forward LSTM and reverse LSTM, and obtaining a merged feature after merging the forward LSTM and the reverse LSTM output, and the merged feature is represented as:
[0087] The forward LSTM is:
[0088] ,
[0089] The reverse LSTM is:
[0090] ,
[0091] The merged feature is:
[0092] ,
[0093] where, and are the outputs of the first i character and the first i- 1character forward LSTM, and are the outputs of the first i character and the first i+ 1character reverse LSTM, is the merged feature;
[0094] The output sequence feature is: ;
[0095] The context semantics of the characters are analyzed by the bidirectional LSTM network, and the bidirectional recurrent network can repair the errors of OCR recognition by using the context.
[0096] S3024, access the enterprise asset database to generate a structured knowledge vector related to the current recognized text , combine the sequence feature and the structured knowledge vector to obtain the final text feature vector representation .
[0097] According to the OCR result, the database is queried to return the most possible equipment information, and the retrieved structured knowledge is converted into a feature vector, which is injected into the model as a strong priori, which can significantly improve the completion ability of the heavily damaged text. For example, the nameplate can only recognize "pressure: 22_V" through the bidirectional LSTM network analysis character, but after the strong guidance of the database priori, it can be inferred as "voltage: 220V".
[0098] S303, determine the fusion weight of the multi-modal feature according to the label damage degree, and combine the fusion weight of the multi-modal feature and the multi-modal feature to obtain the fused multi-modal feature;
[0099] When the nameplate is severely damaged, the visual information may be completely lost, and the text + priori knowledge needs to be relied on; when the characters are stained, the visual texture needs to be relied on.
[0100] In this embodiment, the fusion weight is determined by the image missing degree, and the calculation formula is:
[0101] ,
[0102] where, is the image missing degree score, and the missing degree of the image is obtained according to the damage rate of the segmentation mask, which is obtained by the number of missing pixels / total number of pixels, wherein 0 represents complete, and 1 represents severe missing, wis a weight parameter, b is a bias term, which can be determined by tuning according to specific experimental data;
[0103] Specifically, the fused multi-modal feature is represented as:
[0104] ,
[0105] wherein, represents the fused multi-modal feature, represents the local texture feature and the missing edge perception feature, represents the final text feature vector representation.
[0106] In this embodiment, the basis for dynamic weight adjustment is the degree of image missing. When the image damage degree is greater than the set threshold, it means that the damage is serious, and at this time the visual information is unreliable, and the weight of the visual feature needs to be reduced, and the weight of the text feature including the part based on prior knowledge needs to be increased. When the image damage degree is less than the set threshold, it means that it is relatively complete, and the visual information is reliable, and at this time the weight of the visual feature is increased, and the weight of the text feature is correspondingly reduced or kept balanced.
[0107] When the text recognition result is poor, such as a large area of stains leading to OCR failure, the weight of the visual feature is increased. This feature combines visual details, text context and domain knowledge, and dynamically optimizes the selection of information sources according to the current image state, ensuring that the optimal feature combination can be maintained under different damage conditions.
[0108] S104: obtaining a complete image and text content based on the fused multi-modal feature and the completion reasoning model;
[0109] Specifically, the following steps are included:
[0110] S401: obtaining a character image of image restoration based on the fused multi-modal feature for primary completion;
[0111] The primary completion is based on character morphology analysis, and the missing part is inferred through the stroke continuity of adjacent characters, which is suitable for local small range damage; a non-local mean filter (NLM) algorithm is used, and a preliminary repaired image is output through input of segmented image mask and original image;
[0112] Specifically, the following steps are included:
[0113] S4011: combining multi-modal features, based on character morphology analysis, inferring missing parts through stroke continuity of adjacent characters, and obtaining a morphology repair result;
[0114] The character position decoded from the multi-modal feature is used to locate the bounding box of each character, and each character region is binarized and a skeleton is extracted, the stroke structure of adjacent characters is used to infer the stroke direction of the missing part, and the strokes are drawn in the damaged area according to the known stroke direction and connected with the surrounding strokes;
[0115] S4012, using a non-local mean filter (NLM) algorithm, inputting the segmented image mask and the original image to output a texture repair result;
[0116] For each pixel p to be repaired in the damaged area, a neighborhood window is defined; in the non-damaged area of the image, a similar window is searched, and when calculating the similarity of the two windows, the calculation weights of the pixel value similarity and the feature similarity are considered, and the texture repair result is obtained by weighted average of the pixel value similarity and the feature similarity;
[0117] Wherein, the pixel value similarity can be calculated by Gaussian weighted Euclidean distance, and the feature similarity uses the lightweight ViT feature of the visual branch to calculate the cosine similarity;
[0118] S4013, the morphological repair result and the texture repair result are fused to obtain a primary completion result.
[0119] In the primary completion of multi-modal features, text information (character sequence, position) is used to guide morphological repair to ensure that the stroke structure conforms to the character content, and visual features (local texture) are used to improve NLM to make the repaired texture closer to the texture of the character area (such as sharp stroke edges and uniform interior).
[0120] Through character morphological analysis, it is ensured that the repaired strokes are consistent with the known character structure, for example, when repairing the letter "O", the closed circular structure is maintained.
[0121] S402, based on the primary completion of the character image, a semantic middle-level completion is performed to obtain a semantic completion character image;
[0122] The character image after primary completion may still have missing or blurred areas, and the middle-level completion introduces a Transformer architecture to perform context reasoning by analyzing business knowledge such as asset number verification rules and device model naming rules; specifically including the following steps:
[0123] The character sequence and its corresponding context information are recognized based on the character image after primary completion, the character sequence and its corresponding context information are converted into features suitable for Transformer processing format, the relationship between characters is extracted based on the Transformer encoder, and the final completion suggestion is generated;
[0124] In this embodiment, the features converted into a format suitable for Transformer processing include character embedding vectors and additional context feature vectors;
[0125] Further, when extracting the relationship between characters based on the Transformer encoder, business rule constraints such as asset number verification rules and device model naming rules are added, the input sequence is processed by the Transformer encoder, the complex relationship between characters is captured by the multi-head self-attention mechanism, and the most likely complete character sequence is generated as the final completion suggestion according to the output result of the Transformer.
[0126] Specifically, the features extracted by the business rule constraints are taken as an additional input feature, which is concatenated or added to the output of the Transformer encoder, and then input to the subsequent decoding layer or prediction layer.
[0127] For example, the input device model is CX5?0-SER (one bit missing), the Transformer encoder learns that CX5 often appears in combination with X0, and -SER is a common suffix.
[0128] The corresponding business rules include:
[0129] Rule 1: The 4th digit of the model must be a number (0-9);
[0130] Rule 2: The number following the CX5X0 series represents the memory size and can only be 2, 4, 8, 16;
[0131] Rule 3: The length of the entire model string is fixed at 10 digits;
[0132] The probability distribution of the Transformer predicting the 4th digit (missing bit) may be X (0.1), 0 (0.2), 2 (0.3), 4 (0.25), 8 (0.15);
[0133] Apply rule 1: mask out X -> remaining 0, 2, 4, 8;
[0134] Apply rule 2: further mask out 0 (because CX5X0 can only be followed by 2 / 4 / 8 / 16, but 0 is not in the list, and 16 is two digits, so with one bit missing, it can only be 2 / 4 / 8) -> remaining 2, 4, 8;
[0135] Generate candidate: the model selects the highest probability from 2, 4, 8 (assuming it is 4);
[0136] Final output: CX540-SER, which meets all business rules and utilizes model prediction.
[0137] S403, determine whether the confidence of the recognition result of the character image obtained by the intermediate completion is higher than the set threshold value, if it is lower than the set confidence threshold value, activate the database retrieval mechanism, fuzzy match the recognition result with the asset database, retrieve a plurality of candidate results and sort them, and select the final completion result according to the sorted results.
[0138] In this embodiment, the confidence of the recognition result of the character image obtained by the intermediate completion can be marked by the satisfaction degree of the business rule and the prediction probability. The satisfaction degree of the business rule includes whether the check bit is correct and whether the format is strictly matched. The prediction probability includes the probability of predicting the character at each position in the sequence.
[0139] In this embodiment, the database retrieval mechanism is activated, and the recognition result is fuzzy matched with the asset database, including:
[0140] The text sequence output by the primary and intermediate completion, the context information (such as device type, brand, region, etc.), the asset number format rule (such as length, check bit, coding rule), the enterprise asset database (such as asset ledger, device list, historical record, etc.) are converted into database query statements, and fuzzy matching algorithm (Levenshtein distance) is used for matching to obtain matching results; the matching results are subjected to legal verification and semantic consistency verification to obtain the final matching results;
[0141] In this embodiment, during the legal verification, the check rule of the asset number can be used to verify whether the candidate result is legal;
[0142] During the semantic consistency verification, the format of the asset number can be determined according to the device model; whether the recognized year is reasonable (such as future years may be recognition errors) can be determined;
[0143] It should be noted that the completion strategy can be performed by the edit distance minimum priority method, the semantic rule matching priority method, and the historical frequency priority method.
[0144] The high-level completion activates the database retrieval mechanism, and fuzzy matches part of the recognition result with the enterprise asset library to find the most possible complete information. This hierarchical processing mechanism not only ensures the processing efficiency, but also ensures the completion accuracy. The system can dynamically skip the completed level during operation.
[0145] Figure 3 The asset information intelligent completion system provided by the embodiment of the present application fuses a multi-modal large model, including:
[0146] The image enhancement module 301 is configured to preprocess and enhance the obtained asset nameplate image to obtain an enhanced asset nameplate image.
[0147] an image segmentation module 302, configured to segment the enhanced asset nameplate image to obtain a segmented asset nameplate image;
[0148] a feature extraction module 303, configured to perform multi-modal feature extraction on the segmented asset nameplate image to obtain multi-modal features, dynamically adjust a fusion weight of the multi-modal features according to a missing degree of the image, and obtain fused multi-modal features by combining the fusion weight of the multi-modal features and the multi-modal features;
[0149] an information completion module 304, configured to obtain a complete image and text content based on the fused multi-modal features and a completion inference model.
[0150] It should be noted that the specific implementation of the asset information intelligent completion system of the fusion multi-modal large model according to the embodiments of the present application is similar to the specific implementation of the asset information intelligent completion method of the fusion multi-modal large model according to the embodiments of the present application. For details, please refer to the description in the method part. In order to reduce redundancy, this part will not be repeated here.
[0151] Based on the same inventive concept as the above method, the embodiments of the present application provide a computer device. The computer device can include the asset information intelligent completion system of the fusion multi-modal large model described in the above embodiments. Figure 4 The structure of the computer device in the embodiments of the present application is shown in Figure 4 The computer device 400 can include a processor 401 and a communication interface 402. The communication interface 402 is coupled to the processor 401. The processor 401 obtains the image to be completed through the communication interface 402. The processor 401 is configured to support the computer device 400 to implement the asset information intelligent completion method of the fusion multi-modal large model described in the above embodiments.
[0152] In some possible implementations, as shown by the dashed line in Figure 4 The image retrieval device 400 further includes a memory 403 for storing computer execution instructions and data necessary for the computer device 400. When the computer device 400 is running, the processor 401 executes the computer execution instructions stored in the memory 403, so that the computer device 400 performs the asset information intelligent completion method of the fusion multi-modal large model as described in the above embodiments.
[0153] Based on the same inventive concept as the above method, the embodiments of the present application provide a computer readable storage medium. The computer readable storage medium stores instructions. When the instructions run on the computer, the instructions are used to perform the asset information intelligent completion method of the fusion multi-modal large model as described in the above embodiments.
[0154] The embodiments of the present application further provide a computer program or computer program product, when executed on a computer, cause the computer to implement the asset information intelligent completion method of the fused multi-modal large model as described in the various embodiments above.
[0155] Those skilled in the art will appreciate that the functions described with respect to the various illustrative logical blocks, modules, and algorithm steps described in this specification can be implemented as hardware, software, firmware, or any combination thereof. If implemented in software, the functions described with respect to the various illustrative logical blocks, modules, and steps described in this specification can be stored or transmitted over some computer readable medium as one or more instructions or code for execution by a hardware-based processing unit. Computer-readable media can include computer-readable storage media, which corresponds to a tangible medium such as data storage media, or communication media including any medium that facilitates transfer of the computer program from one place to another, e.g., according to a communication protocol. In this manner, computer-readable media generally can correspond to tangible computer readable storage media that is non-transitory or communication media such as signals or carrier waves. Data storage media can be any available media that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementation of the techniques described in this specification. A computer program product can include a computer-readable medium.
[0156] By way of example, and not limitation, such computer-readable storage media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory, or any other storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any
[0157] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Accordingly, the term "processor," as used herein can refer to any of the foregoing structure or any other structure suitable for implementation of the techniques described herein. In addition, in some aspects, the functions described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated in a combined codec, etc. Also, the techniques could be fully implemented in one or more circuits or logic elements.
[0158] The techniques of this disclosure can be implemented in a wide variety of devices or apparatuses, including a wireless handset, a set top box, a mobile television, a wireless local loop (WLL) device, a computer, a tablet, and so on. The techniques of this disclosure can be implemented in one or more of the following technologies, which are examples and are not meant to be an exhaustive list of possible implementations: application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, micro-controllers, microprocessors, other integrated circuits, and / or discrete circuitry. The term "processor" refers to one or more devices capable of executing a sequence of instructions. The term "logic" refers to the set of operational instructions for
[0159] In the above embodiments, the description of each embodiment is focused on, and the part not described in detail in a certain embodiment can refer to the relevant description of other embodiments.
[0160] The above description is only exemplary specific embodiments of the present application, but the protection scope of the present application is not limited thereto, and any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for intelligent completion of asset information by fusing a multi-modal large model, characterized in that, The method comprises the following steps: The acquired asset nameplate image is preprocessed and image enhanced to obtain an enhanced asset nameplate image; The enhanced asset nameplate image is segmented to obtain a segmented asset nameplate image; Multi-modal feature extraction is performed on the segmented asset nameplate image to obtain multi-modal features, the fusion weight of the multi-modal features is dynamically adjusted according to the missing degree of the image, and the fused multi-modal features are obtained by combining the fusion weight of the multi-modal features and the multi-modal features; The multi-modal feature extraction on the segmented asset nameplate image comprises: Local texture features are extracted based on the segmented asset nameplate image and a visual branch; Sequence features of the recognized characters are analyzed using a bidirectional LSTM network, which comprises: obtaining a character sequence by using OCR recognition, mapping each character to a word embedding vector, inputting the word embedding vector into the bidirectional LSTM network to perform forward LSTM and reverse LSTM, obtaining a combined feature by combining the forward LSTM and the reverse LSTM, and obtaining a final text feature vector representation by fusing the combined feature and prior knowledge embedding; The fusion weight of the multi-modal features is determined by the image missing degree, and the calculation formula is: , wherein, is the image missing degree score, the missing degree of the image is obtained according to the damage rate of the segmentation mask, which is obtained by the number of missing pixels / total number of pixels, wherein 0 represents complete, 1 represents severe missing, w is a weight parameter, and b is a bias term; A complete image and text content are obtained based on the fused multi-modal features and a completion reasoning model, which comprises: Primary completion is performed based on the fused multi-modal features to obtain a primary completion result; Semantic secondary completion is performed based on the character image of the primary completion to obtain a semantically completed character image; It is judged whether the confidence of the recognition result of the semantically completed character image obtained through the secondary completion is higher than a set threshold value, if the confidence is lower than the set confidence threshold value, a database retrieval mechanism is activated, the recognition result is fuzzy matched with an asset database, a plurality of candidate results are retrieved and sorted, and a final completion result is selected according to the sorted result.
2. The asset information intelligent completion method of fusing a multi-modal large model according to claim 1, wherein, The preprocessing and image enhancement of the acquired asset nameplate image comprises: The asset nameplate image is subjected to grayscale processing, and the pixel values of the grayscale image are scaled to adjust the size of the input image; The preprocessed image is subjected to image enhancement, similar oil stains and textures are added to the training image in the aspect of data enhancement, a part of the image is randomly occluded to simulate the case of label missing, different occlusion area ratios and occlusion shapes are set, and finally the enhanced asset nameplate image is obtained.
3. The asset information intelligent completion method of fusing a multi-modal large model according to claim 1, wherein, The improved U-Net network is used for segmenting the enhanced asset nameplate image, including: inputting the enhanced asset nameplate image into an encoder, using a convolution layer and a pooling layer to extract image features, adding an attention mechanism SE before the jump connection at the end of the encoder of the U-Net, inputting the feature map calibrated by the encoder into a decoder, adding a residual block after each upsampling layer of the decoder to reconstruct the features, and setting an edge detection branch at the last layer of the decoder, using the high-dimensional feature map at the last layer of the U-Net decoder as an input variability convolution kernel, learning the offset of the coordinates in the high-dimensional feature map by using parallel convolution, adding one or more residual blocks after the variability convolution, generating an edge mapping representing the probability of each pixel in the image belonging to the edge, and obtaining the final segmentation result.
4. The asset information intelligent completion method of fusing a multi-modal large model according to claim 1, wherein, The preliminary completion result is obtained based on the fused multi-modal features, including: Combining the multi-modal features, based on character morphology analysis, the missing part is inferred by the stroke continuity of adjacent characters to obtain a morphology repair result; Using a non-local mean filtering algorithm, a texture repair result is output by inputting the segmented image mask and the original image; The morphology repair result and the texture repair result are fused to obtain the preliminary completion result.
5. The asset information intelligent completion method of fusing a multi-modal large model according to claim 1, wherein, The semantic mid-level completion of the character image based on the preliminary completion is performed to obtain a semantically completed character image, including: The character sequence and its corresponding context information recognized based on the character image of the preliminary completion; The character sequence and its corresponding context information are converted into features suitable for the processing format of the Transformer, the relationship between characters is extracted based on the Transformer encoder, the business rule constraint is added, the input sequence is deeply processed by the Transformer encoder, the complex relationship between characters is captured by using the multi-head self-attention mechanism, and the most likely complete character sequence is generated as the final completion suggestion.
6. The asset information intelligent completion system of fusing a multi-modal large model, characterized in that, Including: An image enhancement module for pre-processing and image enhancement of the acquired asset nameplate image to obtain an enhanced asset nameplate image; An image segmentation module for segmenting the enhanced asset nameplate image to obtain a segmented asset nameplate image; A feature extraction module for multi-modal feature extraction of the segmented asset nameplate image to obtain multi-modal features, dynamically adjusting the fusion weight of the multi-modal features according to the missing degree of the image, and obtaining fused multi-modal features by combining the fusion weight of the multi-modal features and the multi-modal features; The multi-modal features are obtained by segmenting the asset nameplate image, including: Based on the segmented asset nameplate image and the local texture features extracted by the visual branch; The sequence features of the recognized characters are analyzed using a bidirectional LSTM network, including: obtaining a character sequence by using OCR recognition, mapping each character to a word embedding vector, inputting the word embedding vector into the bidirectional LSTM network for forward LSTM and reverse LSTM, merging the forward LSTM and reverse LSTM outputs to obtain merged features, and fusing the merged features and prior knowledge embeddings to obtain a final text feature vector representation; The fusion weight of the multi-modal feature is determined by the image missing degree, and the calculation formula is: , wherein, is the image missing degree score, the missing degree of the image is obtained according to the damage rate of the segmentation mask, which is obtained by the number of missing pixels / total number of pixels, wherein 0 represents complete, 1 represents severe missing, w is a weight parameter, and b is a bias term; An information completion module is configured to obtain complete image and text content based on the fused multi-modal feature and a completion reasoning model, including: Primary completion is performed based on the fused multi-modal feature to obtain a primary completion result; Semantic middle completion is performed based on the character image of the primary completion to obtain a semantically completed character image; It is determined whether the confidence of the recognition result of the semantically completed character image obtained through the middle completion is higher than a set threshold value, and if it is lower than the set confidence threshold value, a database retrieval mechanism is activated, the recognition result is fuzzy matched with an asset database, a plurality of candidate results are retrieved and sorted, and a final completion result is selected according to the sorted result.
7. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the steps of the asset information intelligent completion method of the fused multi-modal large model according to the program.
Citation Information
Patent Citations
Image enhancement method and system based on multi-prior feature fusion
CN118154440A
Medical image segmentation method based on u-shaped network
WO2022257408A1