Illegal capital image description generation method and system based on multi-modal model
Through the illegal fundraising image description generation method based on multimodal model, combined with ViT, OCR and ViLBERT models, the problem that the existing technology is difficult to fully utilize key elements in the image is solved, and more accurate and detailed image description is achieved, which improves the effect of illegal fundraising risk identification.
Patent Information
- Application Number
- CN202510057759.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-14
- Publication Date
- 2025-06-06
AI Technical Summary
When identifying illegal fundraising activities, it is difficult to fully capture key elements in the image, and fail to fully utilize the correlation between words, icons and scenes in the image, resulting in limited recognition accuracy and comprehensiveness.
An illegal fundraising image description generation method based on multimodal model is adopted, and the text information is extracted through the ViT model and OCR technology are extracted, and the ViLBERT model is used to interact across modal information to generate a joint representation of the fused image and text information, and finally a detailed image description text is generated.
By fully utilizing the visual and text elements in the image, more accurate and detailed image description text is generated, which improves the recognition effect of enterprises' illegal fundraising risks and enhances the high-precision recognition ability of illegal fundraising activities.
Smart Images

Figure CN120107981A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of illegal fund-raising image description generation, and in particular to a method and system for generating illegal fund-raising image description based on a multimodal model. Background Art
[0002] Illegal fund-raising activities have always been one of the major risks to social and economic development due to their concealment and complexity. Illegal fund-raising activities are carried out in a variety of ways, such as publishing false investment information and using Internet platforms to commit fraud. These activities not only seriously infringe on the legitimate rights and interests of investors, but also disrupt the normal financial order. With the popularization of the Internet, the forms and means of illegal fund-raising have become more diverse and complex, and it is increasingly important to study and identify these activities.
[0003] Images play a key role in the spread of illegal fundraising activities. Research on illegal fundraising images can improve the accuracy of identifying corporate illegal fundraising risks. At present, the data sources of illegal fundraising images mainly come from the Internet and social media, including advertisements, chat screenshots, application interfaces, website screenshots, etc. These images contain a lot of potential information that can help identify illegal fundraising activities. For example, through the text, patterns, logos, QR codes and other elements in the image, the activity methods, participants and specific details of illegal fundraising can be revealed.
[0004] At present, illegal fund-raising monitoring mainly relies on the analysis and understanding of relevant text and image data. Image data may contain key text information and visual clues (such as fake authority icons, gold coins, etc.), which are crucial to revealing the nature of illegal fund-raising. However, the field of illegal fund-raising risk identification is basically blank, and the existing general image analysis technology has limitations in processing cross-modal data and cannot fully explore and utilize the correlation information between images and texts.
[0005] There are many challenges in the process of assisting in the identification of corporate illegal fundraising risks based on illegal fundraising image analysis. First, the diversity and complexity of image information increases the difficulty of analysis. Illegal fundraising images may come from different sources, including online advertisements, social media posts, chat screenshots, etc. The format, content and quality of these images vary, requiring a highly flexible and robust analysis method. Secondly, illegal fundraising activities are concealed, and relevant information may be scattered in multiple images and texts, requiring comprehensive analysis and reasoning to reveal hidden illegal behaviors. However, the current mining of hidden information in illegal fundraising images does not fully consider the correlation between text, icons and scenes in the image.
[0006] Mining the hidden information in illegal fundraising images can improve the accuracy of determining whether a company is suspected of illegal fundraising. Existing technologies mainly rely on optical character recognition (OCR) technology to extract text information from images, and then use natural language processing technology to analyze to identify illegal fundraising activities, but this method obviously cannot fully utilize the various information in the image. In addition, OCR technology has not yet been directly applied to the identification of illegal fundraising risks, and there is currently no patented technology specifically for generating descriptions of illegal fundraising images. Existing methods are still blank in the field of illegal fundraising image analysis.
[0007] Therefore, the technical problem that needs to be solved urgently is that the complexity and concealment of illegal fundraising activities make it difficult to fully capture all the key elements in the image by relying solely on text information. Current technology does not make full use of the visual elements in the image and does not fully utilize the relationship between the text, icons and contextual scenes in the image, resulting in limited recognition accuracy and comprehensiveness. Summary of the invention
[0008] The present invention is made to solve the above-mentioned problems, and its purpose is to provide a method and system for generating illegal fund-raising image description based on a multimodal model.
[0009] The present invention provides a method for generating descriptions of illegal fund-raising images based on a multimodal model, which has the following characteristics and specifically includes the following steps: S1, image feature extraction: dividing the image blocks, inputting the image data into the ViT model, and extracting the image feature vector; S2, OCR extraction of text information: using OCR technology to extract text information from illegal fund-raising images, and encoding the text information into high-dimensional features; S3, cross-modal information interaction: inputting text features and image features into the ViLBERT model, performing cross-modal information interaction, and generating a joint representation that integrates the information of the two; and S4, image description generation: generating image description text based on the joint representation to reveal more elements of illegal fund-raising.
[0010] The illegal fund-raising image description generation method based on the multimodal model provided by the present invention may also have the following features: wherein step S1 includes the following sub-steps:
[0011] S1-1, image block division, the original illegal fund-raising image Divide the image into fixed-size blocks, each of which The dimension of is determined by the image's height H, width W, and number of channels C. The image I is divided into N image blocks. Where p×p is the resolution of the image block;
[0012] S1-2, embedding vector generation, each image block is converted into an embedding vector through a linear projection layer, the i-th image block x i Flattened into a one-dimensional vector and transformed into an embedded vector z through the linear projection matrix E i , the process is expressed as:
[0013] z i =E·flatten(x i )+e pos,i #(2) Among them, flatten(x i ) means flattening the i-th image block into a vector, e pos,i It is position embedding, which is used to retain the position information of image blocks;
[0014] S1-3, Transformer encoder processing, all embedded vectors are processed by the Transformer encoder. Each layer of the encoder consists of a multi-layer self-attention mechanism and a feedforward neural network. The self-attention mechanism calculates the query, key, and value matrix. The processing process includes:
[0015] Q i ,K i ,V i =z i-1 W q ,z i-1 W k ,z i-1 W v #(3)
[0016]
[0017] Among them, z i-1 is the embedding vector representation of the output of the previous layer, d k is the dimension of the key, Q, K, V are the query, key, and value matrices respectively, and W q , W k , W v The linear transformation matrices representing queries, keys, and values are used to convert the input embedding vectors into query, key, and value matrices for self-attention calculation. The output of the self-attention layer is normalized and input into the feedforward neural network to generate a high-level feature representation of the image. The processing process is expressed as:
[0018] z′ i =LayerNorm(z i-1 +Attention(Q,K,V))#(5)
[0019] z i =LayerNorm(z′ i +FFN(z′ i))#(6)
[0020] Among them, FFN is a feedforward neural network, LayerNorm is layer normalization, and the image features extracted by the ViT model can be expressed as:
[0021] ImageFeatures=ViT(Image Data)=z#(7)The ViT model combines the above steps to generate a high-dimensional feature representation of the image, preparing data for subsequent multimodal processing.
[0022] The illegal fund-raising image description generation method based on the multimodal model provided by the present invention may also have the following features: wherein step S2 includes the following sub-steps:
[0023] S2-1, image preprocessing, preprocessing the image, including grayscale, binarization and denoising, to improve the accuracy of text recognition. This step ensures that subsequent text area detection and recognition are more accurate by improving image quality;
[0024] S2-2, text area detection, uses Convolutional Neural Networks (CNN) to detect the text area in the image, and obtains the feature map F and the border B of the text area. The specific detection process includes:
[0025] F=CNN(I)#(8)
[0026] B=BoundingBoxDetector(F)#(9)
[0027] Where I is the input image, F is the feature map extracted from the image, and B is the bounding box of the text area detected in the feature map;
[0028] S2-3, text region feature extraction, for each detected text region R, extract the region feature F through the convolution layer R , use Bi-directional Long Short-Term Memory (BiLSTM) for sequence modeling, and generate sequence features S:
[0029] F R =ConvLayers(R)#(10)
[0030] S=BiLSTM(F R )#(11);
[0031] S2-4, text feature decoding, use the CTC (Connectionist Temporal Classification) layer to decode the generated sequence feature S to obtain the text sequence T: T = CTCDecoder (S) # (12);
[0032] S2-5, text feature encoding, uses the BERT model to encode the extracted text sequence T to obtain a high-dimensional feature representation:
[0033] h i =BERT(T i )#(13)
[0034] Among them, T i is the i-th word in the text sequence T, h i is the feature representation of the i-th word in the text sequence. The formula for extracting text information from an image and encoding it using OCR technology can be expressed as:
[0035] Text Features=BERT(OCR(Image Data))=H=[h 1 ,h 2 ,…,h M ]#(14).
[0036] The illegal fund-raising image description generation method based on the multimodal model provided by the present invention may also have the following features: wherein step S3 includes the following sub-steps:
[0037] S3-1, cross-modal attention calculation, the image feature Z extracted by ViT and the text feature H encoded by BERT are input into the ViLBERT model. The ViLBERT model interacts with image and text features through a cross-modal attention mechanism. For the image feature matrix Z and the text feature matrix H, we generate query, key, and value matrices respectively. The calculation process of the cross-modal attention mechanism is as follows:
[0038] Q img =ZW q ,K img =ZW k ,V img =ZW v #(15)
[0039] Q text =HW q ,K text =HW k ,V text =HW v #(16)
[0040]
[0041] Among them, Q img , K img 、V img is the query, key, and value matrix of image features, Q text , K text 、V text is the query, key and value matrix of text features, W q , W k , W v is the projection matrix, which is used to generate the query, key and value matrices, A img-text and A text-img Represent the attention between image and text respectively.
[0042] S3-2, joint representation generation, ViLBERT generates the joint representation J of image and text features as:
[0043] JointRepresentation=ViLBERT(Z,H)=J#(19).
[0044] The illegal fund-raising image description generation method based on the multimodal model provided by the present invention may also have the following features: wherein step S4 includes the following sub-steps:
[0045] S4-1, decoder initialization, takes the joint representation J generated by ViLBERT as input, the initial Transformer decoder, the task of the decoder is to generate description text based on the joint representation.
[0046] S4-2, decode and generate text sequence, the decoder generates a text sequence y at each time step t t The decoding process mainly includes the following two parts: forward propagation: the decoder converts the previously generated text sequence y t and the joint representation J is input to the decoder, y t =Decoder(y <t ,J)#(20), where y <t represents the text sequence generated by the decoder before time step t, y t is the word generated by the decoder at time step t, that is, each time step t will be generated according to the previously generated text sequence y <t and the joint representation J to generate a new word y t Attention mechanism: The decoder processes the input of the current time step through self-attention mechanism and interactive attention mechanism. The self-attention mechanism captures the contextual relationship in the generated sequence, and the interactive attention mechanism combines the joint representation of image and text.
[0047] S4-3, text generation, generates words at each time step by iteration until a complete image description text is generated. The decoder generates new words at each time step based on the previous generation results and the current joint representation J, and finally outputs a complete description text. The image description process can be expressed as:
[0048] Image Caption=Decoder(J)#(21).
[0049] The present invention also provides an illegal fund-raising image description generation system based on a multimodal model, which has the following characteristics, including: an image feature extraction module, which divides the image blocks, inputs the image data into the ViT model, and extracts the image feature vector; an OCR text information extraction module, which uses OCR technology to extract text information from illegal fund-raising images, and encodes the text information into high-dimensional features; a cross-modal information interaction module, which inputs text features and image features into the ViLBERT model, performs cross-modal information interaction, and generates a joint representation that integrates the information of the two; and an image description generation module, which generates image description text based on the joint representation to reveal more illegal fund-raising elements.
[0050] The illegal fund-raising image description generation system based on the multimodal model provided by the present invention may also have the following features: wherein the image feature extraction module is performed according to the following sub-steps:
[0051] S1-1, image block division, the original illegal fund-raising image Divide the image into fixed-size blocks, each of which The dimension of is determined by the image's height H, width W, and number of channels C. The image I is divided into N image blocks. Where p×p is the resolution of the image block;
[0052] S1-2, embedding vector generation, each image block is converted into an embedding vector through a linear projection layer, the i-th image block x i Flattened into a one-dimensional vector and transformed into an embedded vector z through the linear projection matrix E i , the process is expressed as:
[0053] z i =E·flatten(x i )+e pos,i #(2) Among them, flatten(x i ) means flattening the i-th image block into a vector, e pos,i It is position embedding, which is used to retain the position information of image blocks;
[0054] S1-3, Transformer encoder processing, all embedded vectors are processed by the Transformer encoder. Each layer of the encoder consists of a multi-layer self-attention mechanism and a feedforward neural network. The self-attention mechanism calculates the query, key, and value matrix. The processing process includes:
[0055] Q i ,K i ,V i =z i-1 W q ,z i-1 W k ,z i-1 W v #(3)
[0056]
[0057] Among them, z i-1 is the embedding vector representation of the output of the previous layer, d k is the dimension of the key, Q, K, V are the query, key, and value matrices respectively, and W q , W k , W v The linear transformation matrices representing queries, keys, and values are used to convert the input embedding vectors into query, key, and value matrices for self-attention calculation. The output of the self-attention layer is normalized and input into the feedforward neural network to generate a high-level feature representation of the image. The processing process is expressed as:
[0058] z′ i =LayerNorm(z i-1 +Attention(Q,K,V))#(5)
[0059] z i =LayerNorm(z′ i +FFN(z′ i ))#(6)
[0060] Among them, FFN is a feedforward neural network, LayerNorm is layer normalization, and the image features extracted by the ViT model can be expressed as:
[0061] Image Features=ViT(Image Data)=z#(7)The ViT model combines the above steps to generate a high-dimensional feature representation of the image, preparing data for subsequent multimodal processing.
[0062] The illegal fund-raising image description generation system based on the multimodal model provided by the present invention may also have the following features: wherein the OCR text information extraction module is performed according to the following sub-steps: S2-1, image preprocessing, preprocessing the image, including grayscale, binarization and denoising, to improve the accuracy of text recognition. This step ensures that subsequent text area detection and recognition are more accurate by improving image quality;
[0063] S2-2, text area detection, uses a convolutional neural network to detect the text area in the image, obtains the feature map F and the border B of the text area, and the specific detection process includes:
[0064] F=CNN(I)#(8)
[0065] B=BoundingBoxDetector(F)#(9)
[0066] Where I is the input image, F is the feature map extracted from the image, and B is the bounding box of the text area detected in the feature map;
[0067] S2-3, text region feature extraction, for each detected text region R, extract the region feature F through the convolution layer R , use Bi-directional Long Short-Term Memory (BiLSTM) for sequence modeling, and generate sequence features S:
[0068] F R =ConvLayers(R)#(10)
[0069] S=BiLSTM(F R )#(11);
[0070] S2-4, text feature decoding, use the CTC layer to decode the generated sequence feature S to obtain the text sequence T:
[0071] T = CTCDecoder (S) # (12);
[0072] S2-5, text feature encoding, uses the BERT model to encode the extracted text sequence T to obtain a high-dimensional feature representation:
[0073] h i =BERT(T i )#(13) Among them, T i is the i-th word in the text sequence T, h i is the feature representation of the i-th word in the text sequence. The formula for extracting text information from an image and encoding it using OCR technology can be expressed as:
[0074] Text Features=BERT(OCR(Image Data))=H=[h 1 ,h 2 ,…,h M ]#(14).
[0075] In the illegal fundraising image description generation system based on the multimodal model provided by the present invention, it can also have the following features: wherein, the cross-modal information interaction module is performed according to the following sub-steps: S3-1, cross-modal attention calculation, the image feature Z extracted by ViT and the text feature H encoded by BERT are input into the ViLBERT model, and the ViLBERT model interacts with the image and text features through the cross-modal attention mechanism. For the image feature matrix Z and the text feature matrix H, we generate query, key and value matrices respectively, and the calculation process of the cross-modal attention mechanism is as follows:
[0076] Q img =ZW q ,K img =ZW k ,V img =ZW v #(15)
[0077] Q text =HW q ,K text =HW k ,V text =HW v #(16)
[0078]
[0079] Among them, Q img , K img 、V img is the query, key, and value matrix of image features, Q text , K text 、V text is the query, key and value matrix of text features, W q , W k , W v is the projection matrix, which is used to generate the query, key and value matrices, A img-text and A text-img Respectively represent the attention between image and text;
[0080] S3-2, joint representation generation, ViLBERT generates the joint representation J of image and text features as:
[0081] Joint Representation=ViLBERT(Z,H)=J#(19).
[0082] In the illegal fund-raising image description generation system based on the multimodal model provided by the present invention, it can also have the following features: wherein the image description generation module is performed according to the following sub-steps: S4-1, decoder initialization, taking the joint representation J generated by ViLBERT as input, initializing the Transformer decoder, the task of the decoder is to generate a description text according to the joint representation;
[0083] S4-2, decode and generate text sequence, the decoder generates a text sequence y at each time step t t The decoding process mainly includes the following two parts: Forward propagation: The decoder converts the previously generated text sequence y t and the joint representation J is input to the decoder,
[0084] y t =Decoder(y <t ,J)#(20)
[0085] Among them, y <t represents the text sequence generated by the decoder before time step t, y t is the word generated by the decoder at time step t, that is, each time step t will be generated according to the previously generated text sequence y <r and the joint representation J to generate a new word y t ,Attention mechanism: The decoder processes the input of the current time step through the self-attention mechanism and the interactive attention mechanism. The self-attention mechanism captures the contextual relationship in the generated sequence, and the interactive attention mechanism combines the joint representation of image and text;
[0086] S4-3, text generation, generates words at each time step by iteration until a complete image description text is generated. The decoder generates new words at each time step based on the previous generation results and the current joint representation J, and finally outputs a complete description text. The image description process can be expressed as:
[0087] Image Caption=Decoder(J)#(21).
[0088] Functions and Effects of the Invention
[0089] According to the method and system for generating illegal fund-raising image description based on a multimodal model involved in the present invention, the present invention combines ViT (Vision Transformer) to extract image features and OCR technology to extract image text information. Through the cross-modal pre-training model ViLBERT (Vision-and-Language BERT), it can fully utilize various elements in the image, including visual elements and text elements, so as to provide more comprehensive and accurate information support, and then generate accurate image description text, dig out more illegal fund-raising elements hidden in the image, generate more accurate and detailed image description text, and improve the identification effect of illegal fund-raising risks of enterprises.
[0090] Compared with single text or image analysis methods, the method of the present invention has higher accuracy and robustness, can achieve high-precision identification of illegal fund-raising activities in various complex scenarios, improve the identification efficiency and accuracy of illegal fund-raising activities, and provide strong technical support for the supervision of illegal fund-raising activities.
[0091] The method is highly flexible and adaptable, and can adapt to illegal fundraising images of different sources and types, including advertising pictures, social media screenshots, chat records, etc., enhancing the ability to respond to dynamic changes in illegal fundraising activities. In addition, the present invention fills the gap in the generation technology of illegal fundraising image descriptions, helping regulatory agencies to discover and warn of illegal fundraising risks earlier, and has important application value and prospects. It provides new technical means and ideas for the analysis and research of corporate illegal fundraising risks. BRIEF DESCRIPTION OF THE DRAWINGS
[0092] Figure 1 It is a schematic diagram of a method for generating description of an illegal fund-raising image based on OCR, ViT and ViLBERT models in an embodiment of the present invention. DETAILED DESCRIPTION
[0093] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and the accompanying drawings specifically illustrate the illegal fund-raising image description generation method and system based on the multimodal model of the present invention.
[0094] Figure 1 It is a schematic diagram of a method for generating description of an illegal fund-raising image based on OCR, ViT and ViLBERT models in an embodiment of the present invention.
[0095] like Figure 1 As shown, the illegal fund-raising image description generation method based on the multimodal model in this embodiment specifically includes the following steps:
[0096] S1, image feature extraction, divide the image into blocks, input the image data into the ViT model, and extract the image feature vector.
[0097] S1-1, image block division.
[0098] The original illegal fund-raising image Divide the image into fixed-size blocks, each of which The dimension of is determined by the image's height H, width W, and number of channels C. The image I is divided into N image blocks.
[0099]
[0100] Among them, p×p is the resolution of the image block.
[0101] S1-2, embedding vector generation.
[0102] Each image patch is converted into an embedding vector through a linear projection layer, and the i-th image patch x i Flattened into a one-dimensional vector and transformed into an embedded vector z through the linear projection matrix E i , the process is expressed as:
[0103] z i =E·flatten(x i )+e pos,i #(2)
[0104] Among them, flatten(x i ) means flattening the i-th image block into a vector, e pos,i It is a position embedding, which is used to preserve the position information of image blocks.
[0105] S1-3, Transformer encoder processing.
[0106] All embedding vectors are processed through the Transformer encoder. Each layer of the encoder consists of a multi-layer self-attention mechanism and a feed-forward neural network. The self-attention mechanism calculates the query, key, and value matrices. The processing process includes:
[0107] Q i ,K i ,V i =z i-1 W q ,z i-1 W k ,z i-1 W v #(3)
[0108]
[0109] Among them, z i-1 is the embedding vector representation of the output of the previous layer, d kis the dimension of the key, Q, K, V are the query, key, and value matrices respectively, and W q , W k , W v The linear transformation matrices representing query, key, and value, respectively, are used to transform the input embedding vector into query, key, and value matrices for self-attention calculation.
[0110] The output of the self-attention layer is normalized and input into the feedforward neural network to generate a high-level feature representation of the image. The processing process is expressed as:
[0111] z′ i =LayerNorm(z i-1 +Attention(Q,K,V))#(5)
[0112] z i =LayerNorm(z′ i +FFN(z′ i ))#(6)
[0113] Among them, FFN is a feedforward neural network and LayerNorm is layer normalization.
[0114] The image features extracted by the ViT model can be expressed as:
[0115] Image Features=ViT(Image Data)=z#(7)
[0116] The ViT model combines the above steps to generate a high-dimensional feature representation of the image and prepares data for subsequent multimodal processing.
[0117] S2, OCR extracts text information, uses OCR technology to extract text information from illegal fundraising images, and encodes the text information into high-dimensional features.
[0118] S2-1, image preprocessing.
[0119] The image is preprocessed, including grayscale, binarization and denoising, to improve the accuracy of text recognition. This step ensures that subsequent text area detection and recognition are more accurate by improving image quality.
[0120] S2-2, text area detection.
[0121] A convolutional neural network is used to detect the text area in the image to obtain a feature map F and a border B of the text area. The specific detection process includes:
[0122] F=CNN(I)#(8)
[0123] B=BoundingBoxDetector(F)#(9)
[0124] Among them, I is the input image, F is the feature map extracted from the image, and B is the bounding box of the text area detected in the feature map.
[0125] S2-3, text area feature extraction.
[0126] For each detected text region R, the regional features F are extracted through the convolutional layer R , use Bi-directional Long Short-Term Memory (BiLSTM) for sequence modeling, and generate sequence features S:
[0127] F R =ConvLayers(R)#(10)
[0128] S=BiLSTM(F R) #(11).
[0129] S2-4, text feature decoding.
[0130] Use the CTC layer to decode the generated sequence feature S to obtain the text sequence T:
[0131] T=CTCDecoder(S)#(12).
[0132] S2-5, text feature encoding.
[0133] Use the BERT model to encode the extracted text sequence T to obtain a high-dimensional feature representation:
[0134] h i =BERT(T i )#(13)
[0135] Among them, T i is the i-th word in the text sequence T, h i is the feature representation of the i-th word in the text sequence.
[0136] The formula for extracting text information from an image and encoding it using OCR technology can be expressed as:
[0137] Text Features=BERT(OCR(Image Data))=H=[h 1 ,h 2 ,…,h M ]#(14).
[0138] S3, cross-modal information interaction, inputs text features and image features into the ViLBERT model, performs cross-modal information interaction, and generates a joint representation that integrates the information of the two.
[0139] S3-1, cross-modal attention computation.
[0140] The image feature Z extracted by ViT and the text feature H encoded by BERT are input into the ViLBERT model. The ViLBERT model interacts the image and text features through the cross-modal attention mechanism. For the image feature matrix Z and the text feature matrix H, we generate the query, key and value matrices respectively. The calculation process of the cross-modal attention mechanism is as follows:
[0141] Q img =ZW q ,K img =ZW k ,V img =ZW v #(15)
[0142] Q text =HW q ,K text =HW k ,V text =HW v #(16)
[0143]
[0144] Among them, Q img , K img 、V img is the query, key, and value matrix of image features, Q text , K text 、V text is the query, key and value matrix of text features, W q , W k , W v is the projection matrix, which is used to generate the query, key and value matrices, A img-text and A text-img Represent the attention between image and text respectively.
[0145] S3-2, joint representation generation.
[0146] ViLBERT generates a joint representation J of image and text features as:
[0147] Joint Representation=ViLBERT(Z,H)=J#(19).
[0148] S4, image description generation, generates image description text based on the joint representation, revealing more elements of illegal fundraising.
[0149] S4-1, decoder initialization.
[0150] The joint representation J generated by ViLBERT is taken as input to the initial Transformer decoder. The task of the decoder is to generate description text based on the joint representation.
[0151] S4-2, decode and generate text sequence.
[0152] The decoder generates a text sequence y at each time step t t ,The decoding process mainly includes the following two parts.
[0153] Forward propagation: The decoder generates the previously generated text sequence y t The combined representation J is input to the decoder.
[0154] y t =Decoder(y <t ,J)#(20)
[0155] Among them, y <t represents the text sequence generated by the decoder before time step t, y t is the word generated by the decoder at time step t, that is, each time step t will be generated according to the previously generated text sequence y <t and the joint representation J to generate a new word y t .
[0156] Attention mechanism: The decoder processes the input of the current time step through self-attention mechanism and interactive attention mechanism. The self-attention mechanism captures the contextual relationship in the generated sequence, and the interactive attention mechanism combines the joint representation of image and text.
[0157] S4-3, text generation.
[0158] By iteratively generating words at each time step until a complete image description text is generated, the decoder generates new words at each time step based on the previous generation results and the current joint representation J, and finally outputs a complete description text.
[0159] The image description process can be expressed as:
[0160] Image Caption=Decoder(J)#(21).
[0161] Based on the above method, the present invention provides an illegal fundraising image description generation system based on a multimodal model, including an image feature extraction module, an OCR text information extraction module, a cross-modal information interaction module and an image description generation module.
[0162] Image feature extraction module: According to the above step S1, the image is divided into blocks, the image data is input into the ViT model, and the image feature vector is extracted.
[0163] OCR text information extraction module: According to the above step S2, OCR technology is used to extract text information from illegal fund-raising images, and the text information is encoded into high-dimensional features.
[0164] Cross-modal information interaction module: According to the above step S3, the text features and image features are input into the ViLBERT model to perform cross-modal information interaction and generate a joint representation that integrates the information of the two.
[0165] Image description generation module: according to the above step S4, an image description text is generated based on the joint representation to reveal more illegal fund-raising elements.
[0166] Functions and Effects of the Embodiments
[0167] According to the method and system for generating illegal fund-raising image description based on a multimodal model involved in the present invention, the present invention combines ViT (Vision Transformer) to extract image features and OCR technology to extract image text information. Through the cross-modal pre-training model ViLBERT (Vision-and-Language BERT), it can fully utilize various elements in the image, including visual elements and text elements, so as to provide more comprehensive and accurate information support, and then generate accurate image description text, dig out more illegal fund-raising elements hidden in the image, generate more accurate and detailed image description text, and improve the identification effect of illegal fund-raising risks of enterprises.
[0168] Compared with single text or image analysis methods, the method of the present invention has higher accuracy and robustness, can achieve high-precision identification of illegal fund-raising activities in various complex scenarios, improve the identification efficiency and accuracy of illegal fund-raising activities, and provide strong technical support for the supervision of illegal fund-raising activities.
[0169] The method is highly flexible and adaptable, and can adapt to illegal fundraising images of different sources and types, including advertising pictures, social media screenshots, chat records, etc., enhancing the ability to respond to dynamic changes in illegal fundraising activities. In addition, the present invention fills the gap in the generation technology of illegal fundraising image descriptions, helping regulatory agencies to discover and warn of illegal fundraising risks earlier, and has important application value and prospects. It provides new technical means and ideas for the analysis and research of corporate illegal fundraising risks.
[0170] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.
Claims
1. A method for generating descriptions of illegal fund-raising images based on a multimodal model, characterized in that: The specific steps include: S1, image feature extraction: divide the image into blocks, input the image data into the ViT model, and extract the image feature vector; S2, OCR extraction of text information: Use OCR technology to extract text information from illegal fund-raising images and encode the text information into high-dimensional features; S3, cross-modal information interaction: input text features and image features into the ViLBERT model to perform cross-modal information interaction and generate a joint representation that integrates the information of the two; as well as S4, Image description generation: Generate image description text based on joint representation to reveal more elements of illegal fundraising.
2. The method for generating illegal fund-raising image description based on a multimodal model according to claim 1, characterized in that: in, The step S1 includes the following sub-steps: S1-1, image block division, The original illegal fund-raising image Divide the image into fixed-size blocks, each of which The dimension of is determined by the image's height H, width W, and number of channels C. The image I is divided into N image blocks. Where p×p is the resolution of the image block; S1-2, embedding vector generation, Each image patch is converted into an embedding vector through a linear projection layer, and the i-th image patch x i Flattened into a one-dimensional vector and transformed into an embedded vector z through the linear projection matrix E i , the process is expressed as: z i =E·flatten(x i )+e pos,i #(2) Among them, flatten(x i ) means flattening the i-th image block into a vector, e pos,i It is position embedding, which is used to retain the position information of image blocks; S1-3, Transformer encoder processing, All embedding vectors are processed through the Transformer encoder. Each layer of the encoder consists of a multi-layer self-attention mechanism and a feed-forward neural network. The self-attention mechanism calculates the query, key, and value matrices. The processing process includes: Q i ,K i ,V i =z i-1 W q ,z i-1 W k ,z i-1 W v #(3) Among them, z i-1 is the embedding vector representation of the previous layer output, d k is the dimension of the key, Q, K, V are the query, key, and value matrices respectively, and W q , W k , W v The linear transformation matrices representing query, key, and value, respectively, are used to transform the input embedding vector into query, key, and value matrices for self-attention calculation. The output of the self-attention layer is normalized and input into the feedforward neural network to generate a high-level feature representation of the image. The processing process is expressed as: z′ i =LayerNorm(z i-1 +Attention(Q,K,V))#(5) z i =LayerNorm(z′ i +FFN(z′ i ))#(6) Among them, FFN is a feedforward neural network, LayerNorm is layer normalization, The image features extracted by the ViT model can be expressed as: Image Features=ViT(Image Data)=z#(7)The ViT model combines the above steps to generate a high-dimensional feature representation of the image, preparing data for subsequent multimodal processing.
3. According to claim 2, the method for generating illegal fund-raising image description based on a multimodal model, Features: Wherein, step S2 includes the following sub-steps: S2-1, image preprocessing, Preprocess the image, including grayscale, binarization and denoising, to improve the accuracy of text recognition. This step improves the image quality to ensure that subsequent text area detection and recognition are more accurate. S2-2, text area detection, Use a convolutional neural network to detect the text area in the image and obtain the feature map F and the border B of the text area. The specific detection process includes: F=CNN(I)#(8) B=BoundingBoxDetector(F)#(9) Where I is the input image, F is the feature map extracted from the image, and B is the bounding box of the text area detected in the feature map; S2-3, text area feature extraction, For each detected text region R, the regional features F are extracted through the convolutional layer R , use Bi-directional Long Short-Term Memory (BiLSTM) for sequence modeling, and generate sequence features S: F R =ConvLayers(R)#(10) S=BiLSTM(F R )#(11); S2-4, text feature decoding, Use the CTC layer to decode the generated sequence feature S to obtain the text sequence T: T = CTCDecoder (S) # (12); S2-5, text feature encoding, Use the BERT model to encode the extracted text sequence T to obtain a high-dimensional feature representation: h i =BERT(T i )#(13) Among them, T i is the i-th word in the text sequence T, h i is the feature representation of the i-th word in the text sequence, The formula for extracting text information from an image and encoding it using OCR technology can be expressed as: TextFeatures=BERT(OCR(Image Data))=H=[h1,h2,…,h M ]#(14)。 4. The method for generating illegal fund-raising image description based on a multimodal model according to claim 3 is characterized by: in, The step S3 includes the following sub-steps: S3-1, cross-modal attention computation, The image feature Z extracted by ViT and the text feature H encoded by BERT are input into the ViLBERT model. The ViLBERT model interacts the image and text features through the cross-modal attention mechanism. For the image feature matrix Z and the text feature matrix H, we generate the query, key and value matrices respectively. The calculation process of the cross-modal attention mechanism is as follows: Q img =ZW q ,K img =ZW k ,V img =ZW v #(15) Q text =HW q ,K text =HW k ,V text =HW v #(16) Among them, Q img , K img 、V img is the query, key, and value matrix of image features, Q text , K text 、V text is the query, key and value matrix of text features, W q , W k , W v is the projection matrix, which is used to generate the query, key and value matrices, A img-text and A text-img Respectively represent the attention between image and text; S3-2, Joint Representation Generation, ViLBERT generates a joint representation J of image and text features as: Joint Representation=ViLBERT(Z,H)=J#(19).
5. According to claim 4, the method for generating illegal fund-raising image description based on a multimodal model, Features: Wherein, the step S4 includes the following sub-steps: S4-1, decoder initialization, Take the joint representation J generated by ViLBERT as input and the initial Transformer decoder. The task of the decoder is to generate description text based on the joint representation. S4-2, decode and generate text sequence, The decoder generates a text sequence y at each time step t t , the decoding process mainly includes the following two parts: Forward propagation: The decoder generates the previously generated text sequence y t and the joint representation J is input to the decoder, and t =Decoder(and <t ,J)#(20) Among them, y <t represents the text sequence generated by the decoder before time step t, y t is the word generated by the decoder at time step t, that is, each time step t will be generated according to the previously generated text sequence y <t and the joint representation J to generate a new word y t ; Attention mechanism: The decoder processes the input of the current time step through self-attention mechanism and interactive attention mechanism. The self-attention mechanism captures the contextual relationship in the generated sequence, and the interactive attention mechanism combines the joint representation of image and text. S4-3, text generation, By iteratively generating words at each time step until a complete image description text is generated, the decoder generates new words at each time step based on the previous generation results and the current joint representation J, and finally outputs a complete description text. The image description process can be expressed as: Image Caption=Decoder(J)#(21).
6. An illegal fundraising image description generation system based on a multimodal model, characterized in that: include: The image feature extraction module divides the image into blocks, inputs the image data into the ViT model, and extracts the image feature vector; The OCR text information extraction module uses OCR technology to extract text information from illegal fund-raising images and encodes the text information into high-dimensional features; The cross-modal information interaction module inputs text features and image features into the ViLBERT model to perform cross-modal information interaction and generate a joint representation that integrates the information of the two. as well as The image description generation module generates image description text based on the joint representation, revealing more elements of illegal fundraising.
7. The method system for generating illegal fund-raising image description based on multimodal model according to claim 6 is characterized by: in, The image feature extraction module is performed according to the following sub-steps: S1-1, image block division, The original illegal fund-raising image Divide the image into fixed-size blocks, each of which The dimension of is determined by the image's height H, width W, and number of channels C. The image I is divided into N image blocks. Where p×p is the resolution of the image block; S1-2, embedding vector generation, Each image patch is converted into an embedding vector through a linear projection layer, and the i-th image patch x i Flattened into a one-dimensional vector and transformed into an embedded vector z through the linear projection matrix E i , the process is expressed as: z i =E·flatten(x i )+e pos,i #(2) Among them, flatten(x i ) means flattening the i-th image block into a vector, e pos,i It is position embedding, which is used to retain the position information of image blocks; S1-3, Transformer encoder processing, All embedding vectors are processed through the Transformer encoder. Each layer of the encoder consists of a multi-layer self-attention mechanism and a feed-forward neural network. The self-attention mechanism calculates the query, key, and value matrices. The processing process includes: Q i ,K i ,V i =z i-1 W q ,z i-1 W k ,z i-1 W v #(3) Among them, z i-1 is the embedding vector representation of the previous layer output, d k is the dimension of the key, Q, K, V are the query, key, and value matrices respectively, and W q , W k , W v Linear transformation matrices representing query, key, and value, respectively. These matrices are used to transform the input embedding vector into query, key, and value matrices for self-attention calculation; The output of the self-attention layer is normalized and input into the feedforward neural network to generate a high-level feature representation of the image. The processing process is expressed as: z′ i =LayerNorm(z i-1 +Attention(Q,K,V))#(5) z i =LayerNorm(z′ i +FFN(z′ i ))#(6) Among them, FFN is a feedforward neural network, and LayerNorm is layer normalization; The image features extracted by the ViT model can be expressed as: Image Features=ViT(Image Data)=z#(7)The ViT model combines the above steps to generate a high-dimensional feature representation of the image, preparing data for subsequent multimodal processing.
8. According to claim 7, the illegal fundraising image description generation method system based on multimodal model, Features: in, The OCR text information extraction module is performed according to the following sub-steps: S2-1, image preprocessing, Preprocess the image, including grayscale, binarization and denoising, to improve the accuracy of text recognition. This step improves the image quality to ensure that subsequent text area detection and recognition are more accurate. S2-2, text area detection, Use a convolutional neural network to detect the text area in the image and obtain the feature map F and the border B of the text area. The specific detection process includes: F=CNN(I)#(8) B=BoundingBoxDetector(F)#(9) Where I is the input image, F is the feature map extracted from the image, and B is the bounding box of the text area detected in the feature map; S2-3, text area feature extraction, For each detected text region R, the regional features F are extracted through the convolutional layer R , use Bi-directional Long Short-Term Memory (BiLSTM) for sequence modeling, and generate sequence features S: F R =ConvLayers(R)#(10) S=BiLSTM(F R )#(11); S2-4, text feature decoding, Use the CTC layer to decode the generated sequence feature S to obtain the text sequence T: T = CTCDecoder (S) # (12); S2-5, text feature encoding, Use the BERT model to encode the extracted text sequence T to obtain a high-dimensional feature representation: h i =BERT(T i )#(13) Among them, T i is the i-th word in the text sequence T, h i is the feature representation of the i-th word in the text sequence, The formula for extracting text information from an image and encoding it using OCR technology can be expressed as: Text Features=BERT(OCR(Image Data))=H=[h1,h2,…,h M ]#(14)。 9. The illegal fund-raising image description generation method system based on multimodal model according to claim 8 is characterized by: in, The cross-modal information interaction module is performed according to the following sub-steps: S3-1, cross-modal attention computation, The image feature Z extracted by ViT and the text feature H encoded by BERT are input into the ViLBERT model. The ViLBERT model interacts the image and text features through the cross-modal attention mechanism. For the image feature matrix Z and the text feature matrix H, we generate the query, key and value matrices respectively. The calculation process of the cross-modal attention mechanism is as follows: Q img =ZW q ,K img =ZW k ,V img =ZW v #(15) Q text =HW q ,K text =HW k ,V text =HW v #(16) Among them, Q img , K img 、V img is the query, key, and value matrix of image features, Q text , K text 、V text is the query, key and value matrix of text features, W q , W k , W v is the projection matrix, which is used to generate the query, key and value matrices, A img-text and A text-img Respectively represent the attention between image and text; S3-2, Joint Representation Generation, ViLBERT generates a joint representation J of image and text features as: Joint Representation=ViLBERT(Z,H)=J#(19).
10. The illegal fundraising image description generation method system based on multimodal model according to claim 9, Features: in, The image description generation module is performed according to the following sub-steps: S4-1, decoder initialization, Take the joint representation J generated by ViLBERT as input and the initial Transformer decoder. The task of the decoder is to generate description text based on the joint representation. S4-2, decode and generate text sequence, The decoder generates a text sequence y at each time step t t , the decoding process mainly includes the following two parts: Forward propagation: The decoder generates the previously generated text sequence y t and the joint representation J is input to the decoder, and t =Decoder(and <t ,J)#(20) Among them, y <t represents the text sequence generated by the decoder before time step t, y t is the word generated by the decoder at time step t, that is, each time step t will be generated according to the previously generated text sequence y <t and the joint representation J to generate a new word y t ; Attention mechanism: The decoder processes the input of the current time step through self-attention mechanism and interactive attention mechanism. The self-attention mechanism captures the contextual relationship in the generated sequence, and the interactive attention mechanism combines the joint representation of image and text. S4-3, text generation, By iteratively generating words at each time step until a complete image description text is generated, the decoder generates new words at each time step based on the previous generation results and the current joint representation J, and finally outputs a complete description text. The image description process can be expressed as: Image Caption=Decoder(J)#(21).