An image embedding method and system

By acquiring images and metadata in the procuratorial industry, generating text summary using large language models, and combining YOLOv10 and SAM2 models for analysis, the limitations of the multimodal embedding method in the existing technology are solved, and more accurate semantic information reflection and automated analysis are achieved, which improves retrieval efficiency and analysis accuracy.

CN120032378BActive Publication Date: 2025-07-11JIANGXI AXIS COMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510518988.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-11
Estimated Expiration
2045-04-24

AI Technical Summary

Technical Problem

In the prior art, multimodal embedding methods rely on the quality of multimodal models, and the improvement space is limited. The image summary method is not optimized for special needs in industry applications, making it difficult for the embedded semantic information to accurately reflect the significance of the image in a specific field, and the document-related metadata and context information are not fully utilized, which reduces the degree of retrieval efficiency and automation of analysis.

Method used

By obtaining the images and metadata in the prosecution file, using a large language model to generate text summary, and combining the YOLOv10 model and SAM2 model for analysis, the fusion module of Transformer is used for evidence classification and analysis, and quantitative and qualitative analysis results are generated, and finally embedded in the vector database.

Benefits of technology

The depth and accuracy of image analysis are improved, the search efficiency and automation level of analysis are enhanced. The generated analysis results can be used for semantic retrieval and case analysis, improving the search efficiency and automation level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032378B_ABST
    Figure CN120032378B_ABST
Patent Text Reader

Abstract

The present invention provides an image embedding method and system. The method includes obtaining images in a procuratorial case file, preprocessing the images, and then using a large language model to generate a text summary of the preprocessed images; obtaining metadata in the procuratorial case file, and combining the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the procuratorial case file; for each evidence category, inputting the preprocessed image into the YOLOv10 model and the SAM2 model, and using a fusion module based on Transformer for analysis to obtain a quantitative analysis result and a qualitative analysis result; generating a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embedding it into a vector database. Specifically, by inputting the multimodal data of the text summary and the metadata into a specific model for analysis, the depth and accuracy of image analysis are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image embedding, and particularly relates to an image embedding method and system. Background Art

[0002] Both methods of retrieval based on multi-modal embedding and retrieval based on image summarization have limitations when applied to the procuratorial industry. The method of multi-modal embedding depends on the quality of the multi-modal embedding model, and there is limited room for improvement; while the quality of the image summarization method has not been optimized for the special needs in industry applications (such as the evidence category of images, the use classification of images), resulting in the semantic information embedded being difficult to accurately reflect the meaning of images in specific fields.

[0003] In addition, the metadata and context information related to the document are not fully utilized, reducing the practical value of images as industry evidence, and ultimately leading to insufficient understanding of professional terms and specific scenarios in procuratorial business, affecting the accuracy and automation of analysis. Summary of the Invention

[0004] Based on this, an image embedding method and system are provided in the embodiments of the present invention, aiming to solve the problem in the prior art that the retrieval enhancement generation system cannot effectively utilize multi-modal information, resulting in low retrieval efficiency and low automation level of analysis.

[0005] The first aspect of the embodiments of the present invention provides an image embedding method, which is applied to the procuratorial industry scenario. The method includes:

[0006] Obtain the images in the procuratorial file, preprocess the images, and then use a large language model to generate a text summary of the preprocessed images;

[0007] Obtain the metadata in the procuratorial file, and combine the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the procuratorial file;

[0008] For each evidence category, input the preprocessed image into the YOLOv10 model and the SAM2 model, and use a fusion module based on Transformer for analysis to obtain a quantitative analysis result and a qualitative analysis result;

[0009] Generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database.

[0010] Further, the step of combining the text summary and the metadata to perform evidence classification to obtain each evidence category includes:

[0011] Convert the text abstract and the metadata into corresponding vectors;

[0012] Project the vector corresponding to the text abstract and the vector corresponding to the metadata to the same dimension through a linear layer to obtain vectors in the same dimension;

[0013] Calculate the outer product of the two vectors in the same dimension through bilinear pooling and vectorize the outer product to obtain an expanded vector;

[0014] After reducing the dimension of the expanded vector, perform linear projection, or perform sign square root and normalization processing on the result after dimension reduction in sequence to obtain the target vector;

[0015] Input the target vector into a classifier to predict the evidence category.

[0016] Further, the step of converting the text abstract and the metadata into corresponding vectors includes:

[0017] Use a pre-trained text encoder to convert the text abstract into a dense vector;

[0018] Use a pre-trained text encoder to convert the text metadata in the metadata into a first vector;

[0019] After normalizing the numerical metadata in the metadata, embed it into a second vector through a linear layer;

[0020] Convert the evidence type metadata in the metadata into a third vector through an embedding layer;

[0021] Combine the first vector, the second vector, and the third vector to obtain a single vector.

[0022] Further, the step of inputting the preprocessed image into the YOLOv10 model and the SAM2 model for each evidence category and using a Transformer-based fusion module for analysis to obtain quantitative analysis results and qualitative analysis results includes:

[0023] Extract features from the preprocessed image. Among them, in the SAM2 model, use a convolutional neural network to extract features from the segmentation mask of SAM2 to generate a first feature vector, and in the YOLOv10 model, generate a second feature vector;

[0024] Concatenate the first feature vector and the second feature vector to obtain a combined feature vector;

[0025] Process the combined feature vector through a Transformer encoder and finally output a fused representation, where each layer applies a self-attention mechanism;

[0026] Extract quantitative features from the output fusion representation, perform quantitative analysis, and obtain the quantitative analysis result.

[0027] Furthermore, the pre-trained text encoder is one of the BERT text encoder or the CLIP text encoder.

[0028] Furthermore, the calculation formula for applying the self-attention mechanism to each layer is:

[0029]

[0030] where Q, K, and V respectively represent the query, key, and value matrices derived from the combined feature vector, and d k is the dimension of the key.

[0031] Furthermore, in the step of inputting the preprocessed image into the YOLOv10 model and the SAM2 model and using the Transformer-based fusion module for analysis, the Transformer-based fusion error expression is:

[0032]

[0033] where represents the Transformer-based fusion error, represents the fusion representation, represents the true feature representation.

[0034] The second aspect of the embodiments of the present invention provides an image embedding system for implementing the image embedding method described in the first aspect. The system includes:

[0035] An acquisition module for acquiring images in the prosecution file, preprocessing the images, and then generating a text summary of the preprocessed images using a large language model;

[0036] A classification module for acquiring the metadata in the prosecution file, and combining the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the prosecution file;

[0037] An analysis module for inputting the preprocessed images into the YOLOv10 model and the SAM2 model for each of the evidence categories, and using the Transformer-based fusion module for analysis to obtain a quantitative analysis result and a qualitative analysis result;

[0038] A generation module for generating a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embedding it into a vector database.

[0039] In a third aspect of the embodiments of the present invention, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the image embedding method provided in the first aspect is implemented.

[0040] In a fourth aspect of the embodiments of the present invention, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor, and when the processor executes the program, the image embedding method provided in the first aspect is implemented.

[0041] An image embedding method and system provided in the embodiments of the present invention. The method includes obtaining images in a procuratorial case file, preprocessing the images, and then using a large language model to generate a text summary of the preprocessed images; obtaining metadata in the procuratorial case file, and combining the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the procuratorial case file; for each evidence category, inputting the preprocessed image into the YOLOv10 model and the SAM2 model, and using a fusion module based on Transformer for analysis to obtain a quantitative analysis result and a qualitative analysis result; generating a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embedding it into a vector database. Specifically, by inputting the multimodal data of the text summary and the metadata into a specific model for analysis, the depth and accuracy of image analysis are effectively improved. In addition, the obtained analysis results can be directly used for semantic retrieval, case analysis, and report generation, improving the retrieval efficiency and the automation level of analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1 It is a flowchart of an image embedding method provided in Embodiment 1 of the present invention;

[0043] Figure 2 It is a structural block diagram of an image embedding system provided in Embodiment 2 of the present invention;

[0044] Figure 3 It is a structural block diagram of an electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0045] To facilitate the understanding of the present invention, the present invention will be described more comprehensively below with reference to the relevant drawings. Several embodiments of the present invention are given in the drawings. However, the present invention can be implemented in many different forms and is not limited to the embodiments described herein. On the contrary, these embodiments are provided to make the disclosure of the present invention more thorough and comprehensive.

[0046] It should be noted that when an element is referred to as being "fixed to" another element, it can be directly on the other element or there can also be an intermediate element. When an element is considered to be "connected" to another element, it can be directly connected to the other element or there may be an intermediate element at the same time. The terms "vertical", "horizontal", "left", "right" and similar expressions used herein are for illustrative purposes only.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the technical field to which this invention belongs. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items.

[0048] Embodiment 1

[0049] According to an embodiment of the present invention, there is provided an embodiment of an image embedding method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0050] In this first embodiment, an image embedding method is provided, which can be used in an electronic device, such as a computer. Please refer to Figure 1 , Figure 1 which shows an implementation flowchart of an image embedding method provided in the first embodiment of the present invention, specifically including steps S01 to S04.

[0051] Step S01, obtain the image in the prosecution file, preprocess the image, and then use a large language model to generate a text summary of the preprocessed image.

[0052] In this embodiment, the preprocessing of the image includes: 1) verifying the image format and quality; 2) resizing or normalizing the image to meet the input requirements of the model; 3) using preprocessing techniques to remove noise or artifacts or correct skewness.

[0053] After preprocessing, use large language models such as BLIP, MiniGPT-4 or CLIP to generate a descriptive caption or summary, that is, a text summary, for the input image. For example, "a damaged car with broken glass and brake marks around it".

[0054] Step S02, obtain the metadata in the prosecution file, and combine the text summary and the metadata to perform evidence classification to obtain each evidence category.

[0055] Among them, the metadata is information related to the case extracted from the prosecution file. For example, "Charge: Traffic accident" (text), "Case number: 123456" (number), or "Evidence type: Physical evidence" (evidence type).

[0056] Specifically, converting the text summary and the metadata into corresponding vectors. It should be noted that using a pre-trained text encoder, such as the text encoder of BERT or CLIP, to convert the text summary into a dense vector. Among them, the text summary embedding representation corresponding to the image is: , where d1 is the dimension of the text summary embedding corresponding to the image (for example, 768 for BERT and 512 for CLIP). It can be understood that for "Damaged car, with broken glass and brake marks around", x may be a 512-dimensional vector to capture its semantics;

[0057] Using the same pre-trained text encoder as the one used to convert the text summary into a vector, convert the text metadata in the metadata into a first vector, denoted as: , for example, "Traffic accident" is converted into a 512-dimensional vector;

[0058] After normalizing the numerical metadata in the metadata, for example, "Case number: 123456" is normalized to 0.123456 and embedded into a second vector through a linear layer, that is (for example ):

[0059] Among them, represents the second vector, represents the weight matrix, represents the bias vector, n is the normalized number, , ;

[0060] Convert the evidence type metadata in the metadata into a third vector through an embedding layer, that is , represents the third vector, for example, m = 10;

[0061] Combine the first vector, the second vector, and the third vector to obtain a single vector, that is

[0062] Among them, y represents the single vector, , is the dimension of the metadata embedding. It can be understood that, that is , from the above steps, it can be seen that the metadata "Traffic accident, Case number: 123456, Physical evidence" can be converted into ;

[0063] Project the vector corresponding to the text summary and the vector corresponding to the metadata to the same dimension through a linear layer to obtain vectors in the same dimension. It should be noted that if , project the two vectors to the same dimension d through a linear layer, which is expressed as:

[0064]

[0065] Among them, and both represent weight matrices, and both represent bias vectors, , , d is the selected target dimension (e.g., 512), , is the encoded text summary vector, , is the encoded metadata vector;

[0066] Calculate the outer product of the two vectors in the same dimension through bilinear pooling and vectorize the outer product to obtain the expanded vector. It can be understood that the purpose of calculating the outer product is to capture their pairwise interactions, which is expressed as:

[0067]

[0068] Among them, , is the resulting bilinear matrix (e.g., ), here, each element represents the interaction between the i-th feature of the text summary corresponding to the image and the j-th feature of the metadata. For example, "damaged car" in the text summary corresponding to the image may have a strong interaction with "traffic accident" in the metadata.

[0069] Furthermore, flatten Z into a vector z:

[0070]

[0071] Among them, , is a high-dimensional vector;

[0072] After dimensionality reduction processing of the expanded vector, perform linear projection, or perform sign square root and normalization processing on the result of dimensionality reduction processing in sequence to obtain the target vector. The linear projection is expressed as:

[0073] Among them, represents the weight matrix used in the projection, represents the bias vector in the projection, , , where k is the target dimension (e.g., 1024 or 2048), and the symbol square root and normalization processes are performed in sequence and expressed as:

[0074]

[0075] The aim is to stabilize training and reduce overfitting. z1 is the vector after dimensionality reduction, is the target vector, is the symbol square root result;

[0076] Input the target vector into a classifier to predict the evidence category. Among them, the classifier can be a multi-layer perceptron (MLP), a support vector machine (SVM).

[0077] Step S03, for each of the evidence categories, input the preprocessed image into the YOLOv10 model and the SAM2 model, and use a Transformer-based fusion module for analysis to obtain quantitative analysis results and qualitative analysis results.

[0078] Specifically, the YOLOv10 model is used to train for different evidence categories (e.g., vehicle collisions, forensic analysis, property damage, etc.), adjust the anchor boxes to fit the target shape, and optimize the detection performance; the SAM2 model is used to fine-tune using the same dataset to enhance the segmentation ability for complex regions (such as damaged parts).

[0079] It should be noted that feature extraction is performed on the preprocessed image. Among them, in the SAM2 model, for the segmentation mask of SAM2 Use a convolutional neural network (CNN) to extract features and generate the first feature vector , and after flattening, the size is (such as D = 64, d = 16). In the YOLOv10 model, generate the second feature vector , , and the size is 4 + k;

[0080] Concatenate the first feature vector and the second feature vector to obtain a combined feature vector, expressed as , and the size is ;

[0081] Process the combined feature vector through a Transformer encoder, and finally output the fused representation , where the self-attention mechanism is applied to each layer, and the calculation formula is:

[0082]

[0083] Among them, Q, K, and V respectively represent the query, key, and value matrices derived from the combined feature vector, and d k is the dimension of the key. Among them, the weights for detecting and segmenting the output are adaptively learned by the Transformer to improve the accuracy. The fusion error expression based on the Transformer is:

[0084]

[0085] Among them, represents the fusion error based on the Transformer, represents the fusion representation, represents the true feature representation;

[0086] Quantitative features are extracted from the output fusion representation for quantitative analysis to obtain the quantitative analysis result. Taking a traffic accident as an example, quantitative features are extracted from :

[0087] Area:

[0088]

[0089] Severity:

[0090] Among them, and are preset thresholds (such as pixels,[[]] pixels), A is the area, is the value of the pixel (i, j) on the segmentation mask M output by the SAM2 model (1 for belonging to the target and 0 otherwise); Enhancement analysis: Use a fully connected layer to predict the severity from :

[0091]

[0092] Example: The fused mask shows that the damaged area is 5000 pixels and the prediction is "moderate damage".

[0093] In some other embodiments of the present invention, the prediction bounding box adjustment factor is obtained from :

[0094] Among them, . This feedback loop optimizes the positioning accuracy and is applicable to complex scenarios. Further, feature extraction based on the fusion representation does not require manual rules, such as area:

[0095] Combining the class probability to enhance the severity assessment:

[0096]

[0097] wherein is the balance factor (such as ).

[0098] In addition, by performing image recognition on the preprocessed image, qualitative analysis can be completed. Exemplarily, "The front part of the white sedan is severely damaged: the front of the car is dented, the windshield is broken, the front bumper has fallen off, and the left front tire is ruptured and misaligned, which conforms to the characteristics of a high-speed frontal collision."

[0099] Step S04, generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database.

[0100] Among them, the fused bounding box and segmentation mask are superimposed on the original image to generate an intuitive analysis report. In addition, the generation of the automated report can use a text generation tool (such as the GPT model) fine-tuned with a domain-specific language.

[0101] In summary, the image embedding method in the above embodiments of the present invention obtains images in the procuratorial case files, preprocesses the images, and then uses a large language model to generate a text summary of the preprocessed images; obtains metadata in the procuratorial case files, combines the text summary and the metadata, and performs evidence classification to obtain each evidence category, wherein the metadata is information related to the case extracted from the procuratorial case files; for each evidence category, input the preprocessed image into the YOLOv10 model and the SAM2 model, and use a fusion module based on Transformer for analysis to obtain quantitative analysis results and qualitative analysis results; generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database. Specifically, by inputting the multimodal data of the text summary and the metadata into a specific model for analysis, the depth and accuracy of image analysis are effectively improved. In addition, the obtained analysis results can be directly used for semantic retrieval, case analysis, and report generation, improving the retrieval efficiency and the automation level of analysis.

[0102] Embodiment 2

[0103] Please refer to Figure 2 , Figure 2 which is a structural block diagram of an image embedding system provided by Embodiment 2 of the present invention. The image embedding system 200 is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0104] Specifically, the image embedding system 200 includes: an acquisition module 21, a classification module 22, an analysis module 23, and a generation module 24, where:

[0105] The acquisition module 21 is configured to acquire images in the prosecution file, preprocess the images, and then use a large language model to generate a text summary of the preprocessed images;

[0106] The classification module 22 is configured to acquire the metadata in the prosecution file, and perform evidence classification in combination with the text summary and the metadata to obtain each evidence category, where the metadata is information related to the case extracted from the prosecution file;

[0107] The analysis module 23 is configured to input the preprocessed images into the YOLOv10 model and the SAM2 model for each evidence category, and use a Transformer-based fusion module for analysis to obtain a quantitative analysis result and a qualitative analysis result;

[0108] The generation module 24 is configured to generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database.

[0109] Further, in some alternative embodiments of the present invention, the classification module 22 includes:

[0110] A conversion unit configured to convert the text summary and the metadata into corresponding vectors;

[0111] A projection unit configured to project the vector corresponding to the text summary and the vector corresponding to the metadata to the same dimension through a linear layer to obtain vectors in the same dimension;

[0112] A calculation unit configured to calculate the outer product of two vectors in the same dimension through bilinear pooling and perform outer product vectorization to obtain an expanded vector;

[0113] A first processing unit configured to perform linear projection after dimensionality reduction processing on the expanded vector, or perform sign square root and normalization processing on the result of the dimensionality reduction processing in sequence to obtain a target vector;

[0114] An input unit configured to input the target vector into a classifier to predict the evidence category.

[0115] Further, in some alternative embodiments of the present invention, the conversion unit includes:

[0116] The first transformation subunit is used to transform the text summary into a dense vector using a pre-trained text encoder, and the pre-trained text encoder is one of the BERT text encoder or the CLIP text encoder;

[0117] The second transformation subunit is used to transform the text metadata in the metadata into a first vector using a pre-trained text encoder;

[0118] The normalization processing subunit is used to perform normalization processing on the numerical metadata in the metadata and then embed it into a second vector through a linear layer;

[0119] The third transformation subunit is used to transform the evidence type metadata in the metadata into a third vector through an embedding layer;

[0120] The combination subunit is used to combine the first vector, the second vector, and the third vector to obtain a single vector.

[0121] Furthermore, in some optional embodiments of the present invention, the analysis module 23 includes:

[0122] The feature extraction unit is used to extract features from the preprocessed image. Among them, in the SAM2 model, the segmentation mask of SAM2 is used to extract features using a convolutional neural network to generate a first feature vector, and in the YOLOv10 model, a second feature vector is generated;

[0123] The splicing unit is used to splice the first feature vector and the second feature vector to obtain a combined feature vector;

[0124] The second processing unit is used to process the combined feature vector through a Transformer encoder and finally output a fused representation. Among them, the self-attention mechanism is applied to each layer, and the calculation formula is:

[0125]

[0126] Among them, Q, K, and V respectively represent the query, key, and value matrices derived from the combined feature vector, and d k is the dimension of the key. The fusion error expression based on Transformer is:

[0127]

[0128] Among them, represents the fusion error based on Transformer, represents the fused representation, represents the true feature representation;

[0129] An analysis unit is configured to extract quantitative features from the output fusion representation, perform quantitative analysis, and obtain the quantitative analysis result.

[0130] Embodiment III

[0131] On the other hand, the present invention further provides an electronic device. Please refer to Figure 3 , which shows the electronic device in Embodiment III of the present invention, including a memory 20, a processor 10, and a computer program 30 stored in the memory and executable on the processor. When the processor 10 executes the computer program 30, the image embedding method as described above is implemented.

[0132] Among them, in some embodiments, the processor 10 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips, and is used to run the program code stored in the memory 20 or process data, such as executing an access restriction program, etc.

[0133] Among them, the memory 20 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a magnetic disk, an optical disk, etc. The memory 20 may be an internal storage unit of the electronic device in some embodiments, such as the hard disk of the electronic device. The memory 20 may also be an external storage device of the electronic device in other embodiments, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. Further, the memory 20 may also include both an internal storage unit and an external storage device of the electronic device. The memory 20 can be used not only to store application software and various types of data of the electronic device, but also to temporarily store data that has been output or will be output.

[0134] It should be noted that Figure 3 the structure shown does not limit the electronic device. In other embodiments, the electronic device may include fewer or more components than shown in the figure, or combine some components, or have a different component layout.

[0135] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the image embedding method as described above is implemented.

[0136] Those skilled in the art will understand that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0137] More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection part (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0138] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having suitable combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0139] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0140] The above embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the scope of the patent for the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention. Therefore, the protection scope of the patent for the present invention shall be subject to the appended claims.

Claims

1. An image embedding method, characterized in that, Applied to the procuratorial industry scenario, the method includes: Obtain images in the procuratorial case file, preprocess the images, and then use a large language model to generate a text summary of the preprocessed images; Obtain the metadata in the procuratorial case file, and combine the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the procuratorial case file; For each evidence category, input the preprocessed image into the YOLOv10 model and the SAM2 model, and use a Transformer-based fusion module for analysis to obtain a quantitative analysis result and a qualitative analysis result; Generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database; The step of inputting the preprocessed image into the YOLOv10 model and the SAM2 model for each evidence category and using a Transformer-based fusion module for analysis to obtain a quantitative analysis result and a qualitative analysis result includes: Extract features from the preprocessed image. Among them, in the SAM2 model, use a convolutional neural network to extract features from the segmentation mask of SAM2 to generate a first feature vector, and in the YOLOv10 model, generate a second feature vector; Concatenate the first feature vector and the second feature vector to obtain a combined feature vector; Process the combined feature vector through a Transformer encoder, and finally output a fused representation, where each layer applies a self-attention mechanism; Extract quantitative features from the output fused representation for quantitative analysis to obtain the quantitative analysis result.

2. The image embedding method according to claim 1, wherein The step of combining the text summary and the metadata to perform evidence classification to obtain each evidence category includes: Convert the text summary and the metadata into corresponding vectors; Through a linear layer, project the vector corresponding to the text summary and the vector corresponding to the metadata to the same dimension to obtain vectors in the same dimension; Calculate the outer product of the two vectors in the same dimension through bilinear pooling and perform outer product vectorization to obtain an expanded vector; After dimensionality reduction processing of the expanded vector, perform linear projection, or sequentially perform sign square root and normalization processing on the result of dimensionality reduction processing to obtain a target vector; Input the target vector into a classifier to predict the evidence category.

3. The image embedding method according to claim 2, wherein The step of converting the text summary and the metadata into corresponding vectors includes: Use a pre-trained text encoder to convert the text summary into a dense vector; Use a pre-trained text encoder to convert the text metadata in the metadata into a first vector; After normalizing the numerical metadata in the metadata, embed it into a second vector through a linear layer; Convert the evidence type metadata in the metadata into a third vector through an embedding layer; Combine the first vector, the second vector, and the third vector to obtain a single vector.

4. The image embedding method according to claim 3, wherein The pre-trained text encoder is one of the BERT text encoder or the CLIP text encoder.

5. The image embedding method according to claim 4, wherein, The calculation formula for applying the self-attention mechanism to each layer is as follows: Among them, Q, K, and V respectively represent the query, key, and value matrices derived from the combined feature vector, and d k is the dimension of the key.

6. The image embedding method according to claim 5, wherein In the step of inputting the preprocessed image into the YOLOv10 model and the SAM2 model and using the Transformer-based fusion module for analysis, the expression of the fusion error based on Transformer is: Among them, represents the fusion error based on Transformer, represents the said fusion representation, represents the true feature representation.

7. An image embedding system, characterized in that, For implementing the image embedding method according to any one of claims 1-6, the system includes: An acquisition module, configured to acquire images in the prosecution file, preprocess the images, and then use a large language model to generate a text summary of the preprocessed images; A classification module, configured to acquire the metadata in the prosecution file, and combine the text summary and the metadata to perform evidence classification to obtain each evidence category, where the metadata is information related to the case extracted from the prosecution file; An analysis module, configured to, for each of the evidence categories, input the preprocessed image into the YOLOv10 model and the SAM2 model, and use the Transformer-based fusion module for analysis to obtain a quantitative analysis result and a qualitative analysis result; A generation module, configured to generate a report according to the text summary, the quantitative analysis result, and the qualitative analysis result, and embed it into the vector database; The analysis module includes: A feature extraction unit, configured to extract features from the preprocessed image. Among them, in the SAM2 model, convolutional neural network is used to extract features from the segmentation mask of SAM2 to generate a first feature vector, and in the YOLOv10 model, a second feature vector is generated; A splicing unit, configured to splice the first feature vector and the second feature vector to obtain a combined feature vector; A second processing unit, configured to process the combined feature vector through a Transformer encoder and finally output a fusion representation, where the self-attention mechanism is applied to each layer; An analysis unit, configured to extract quantitative features from the output fusion representation to perform quantitative analysis to obtain the quantitative analysis result.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the image embedding method according to any one of claims 1-6.

9. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the program, it implements the image embedding method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Docket analysis methods and systems

    AU2020100413A4

  • Image electronic evidence screening method based on deep learning

    CN112464015A