Data analysis problem generation method based on image input and large model combination

By building a dual-stream input architecture for images and text and semantic consistency detection, the difficult problem of converting image data into natural language output is solved, and efficient and standardized data analysis problem generation is achieved to meet the needs of professional scenarios.

CN120632138APending Publication Date: 2025-09-12HUAZHONG NORMAL UNIV
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202510683138.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively convert complex data in images into natural language output that meets task requirements. Especially in the generation of data analysis questions, the model lacks a deep understanding of image information, logical consistency, and semantic integrity, and the prompt words lack adaptive capabilities, resulting in chaotic generated content structure and ambiguous meaning.

Method used

A dual-stream input architecture of images and texts based on the visual language model and the large language model of the Qwen-VL architecture is constructed. Learnable prompt word embedding, supervised fine-tuning and semantic consistency verification strategies are introduced. Through image preprocessing, visual feature extraction, structured text generation and semantic consistency detection, efficient and standardized data analysis question generation is achieved.

Benefits of technology

The accuracy and stability of image semantic modeling have been improved. The generated questions have strong logical integrity, standardized terminology, natural question expression, strong ability to adapt to professional scenarios, and improved automation level and generation efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632138A_ABST
    Figure CN120632138A_ABST
Patent Text Reader

Abstract

The invention discloses a data analysis problem generation method based on image input and large model combination, and the method comprises the following steps: S1, obtaining original image data, and carrying out the preprocessing of the original image data; s2, constructing a visual language model based on a Qwen-VL architecture, and performing bidirectional alignment to generate a visual feature vector; s3, constructing a cue word template, and fusing the cue word template through a cross attention mechanism to generate structured text description; s4, constructing a large language model based on a Transform architecture, and performing supervision and fine tuning by adopting a LoRA method to generate a data analysis problem candidate sequence; s5, semantic consistency detection and structural rule matching are carried out, and sequences which do not meet semantic specifications or structural constraints are removed; and S6, constructing an image-text alignment triple, and writing the image-text alignment triple into the data structure in the JSON format for coding and storage. According to the method, the image can be converted into a data analysis problem, and the text generation quality and efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of human-computer interaction and intelligent information processing technology, and in particular to a method for generating data analysis problems based on a combination of image input and a large model. Background Art

[0002] Against the backdrop of the rapid development of artificial intelligence, the ability to process image and text information has become an important research direction in multimodal learning. Especially as Visual Language Models (VLMs) and Large Language Models (LLMs) have gradually matured, how to effectively convert the information in images into text descriptions with semantic levels, and then generate natural language output that meets task requirements, has become a research hotspot in many fields such as education, business, and government affairs. For example, in the field of educational evaluation, tasks such as test question generation and data analysis increasingly rely on the ability to understand complex image data and organize language. Traditional rule-based or manual template-based systems are often unable to adapt to the processing needs of large-scale, unstructured image data, nor can they achieve high-quality automated question generation.

[0003] Visual language models can map image and text information into a unified semantic space, enabling image content understanding and text generation. However, the generalization ability and adaptability of such models to professional data in practice are still limited. In particular, when faced with complex numerical information, graphical logic, and mixed text in images, the models often cannot accurately extract key data points and their inherent relationships, resulting in over-generalization of generated text or omission of core content. In addition, most current VLMs lack targeted optimization of task objectives, especially in generation tasks such as "data analysis problems" that require multi-step reasoning and structural specification control. It is difficult to ensure the logical consistency, semantic integrity, and format compliance of the generated content.

[0004] At the same time, large language models have demonstrated excellent performance in various natural language processing tasks in recent years. Their language modeling capabilities, acquired through training on large-scale corpora, give them significant advantages in text generation, semantic understanding, and logical reasoning. However, LLMs are essentially unimodal language models and lack the ability to directly process image information. If they are used directly for image-driven question generation tasks, they must rely on external components or modules to pre-process the image content. Some current work attempts to use OCR to identify the text content in images and then input the text into the LLM for processing. However, this method lacks modeling of the overall structure of the image, the relationship between data, and the cross-semantics of images and text, and cannot fully utilize the rich contextual information carried by the image. Therefore, it is difficult to generate question text with professional depth and expression specifications by relying solely on the OCR+LLM approach.

[0005] Furthermore, in terms of prompt word design, although Prompt Engineering technology has been widely used to guide large models to complete various tasks, most prompt words are still statically constructed and lack context-awareness and adaptive mechanisms. When processing professional image data, if prompt words cannot flexibly adjust their semantic focus or format control, the accuracy and relevance of the model output will be greatly affected. In multimodal systems, prompt words also need to be reasonably integrated with image features, which requires the prompt words themselves to have certain semantic modeling capabilities and task awareness capabilities. Current mainstream VLMs or LLMs have not yet formed a unified, dynamic, and structure-enhanced prompt word generation mechanism, nor do they lack prompt template optimization methods that are jointly trained with image features.

[0006] To address these issues, some research has begun attempting to integrate multimodal models with language models to enable cross-modal information processing and reasoning. Some work has constructed joint image-text representation networks, attempting to extract structured information from images before generating text. However, these models are mostly focused on general scenarios, such as image captioning or visual question answering (VQA), and their adaptability to specialized domains remains insufficient. For example, in the "data analysis" scenario, models are often required to identify data trends, extreme values, proportional relationships, text annotations, and chart structures within images. These contents are not only highly information-dense but also highly logically organized. When faced with such images, existing models not only struggle to fully extract valuable analytical elements, but the generated questions are often disorganized and ambiguous, failing to meet the needs of practical applications such as educational testing and government report analysis.

[0007] Furthermore, the fusion mechanism between the visual encoder and language decoder in current multimodal model architectures presents numerous bottlenecks. In particular, the lack of fine-grained attention control and adjustable gating mechanisms when image features interact with textual cues leads to unstable information flow in the fused features, making the generated content prone to problems such as loss of image information and disconnected text representation. Furthermore, low image-text alignment accuracy reduces the semantic modeling efficiency of subsequent large language models, impacting the quality of the resulting generated questions.

[0008] Therefore, how to provide a data analysis problem generation method based on the combination of image input and large models is an issue that those skilled in the art urgently need to solve. Summary of the Invention

[0009] One purpose of the present invention is to propose a data analysis question generation method based on the combination of image input and a large model. The present invention makes full use of the image understanding ability of the visual language model and the text generation and logical reasoning ability of the large language model, constructs a dual-stream input architecture of images and text and a prompt word fusion mechanism, and introduces learnable prompt word embedding, supervised fine-tuning and semantic consistency verification strategies. It describes in detail the entire process from image preprocessing, visual feature extraction, structured text generation to data analysis question output, and has the advantages of high degree of automation in question generation, accurate semantic expression, strong structural standardization and strong ability to adapt to professional scenarios.

[0010] According to an embodiment of the present invention, a method for generating data analysis questions based on a combination of image input and a large model includes the following steps:

[0011] S1. Obtain original image data and perform preprocessing;

[0012] S2. Build a visual language model based on the Qwen-VL architecture, input the preprocessed image data into the visual encoder, extract image features and text semantics, and perform bidirectional alignment to generate a visual feature vector;

[0013] S3. Construct a prompt word template, input the prompt word into the visual language model through fine-tuning, and fuse it with the visual feature vector through a cross-attention mechanism to generate a structured text description containing numerical information, data relationships and text elements;

[0014] S4. Build a large language model based on the Transformer architecture and use the LoRA method for supervised fine-tuning. Take structured text descriptions as input sequences and generate candidate sequences of data analysis questions.

[0015] S5. Perform semantic consistency check and structural rule matching on candidate sequences of data analysis questions, eliminate sequences that do not meet semantic specifications or structural constraints, and output the data analysis question sequences that pass the check;

[0016] S6. Combine the original image data, structured text description, and data analysis question sequence into an image-text alignment triplet, and write it into a JSON format data structure for encoding and storage.

[0017] Optionally, the original image data consists of visual data content, including tables, graphics, statistical figures, coordinate axis labels and text descriptions.

[0018] Optionally, the preprocessing includes size normalization, color channel standardization and format conversion.

[0019] Optionally, the Qwen-VL architecture adopts a dual-stream input structure of images and text, and the visual encoder includes two parts: image encoding and text encoding. In the image encoding part, image features under different receptive fields are extracted, and image feature vectors are generated through residual connection, layer normalization and linear mapping. In the text encoding part, position encoding and self-attention mechanism are used to obtain text semantics, generate text embedding representation, and perform bidirectional alignment of the image feature vector and the text embedding representation in a multi-head cross-attention layer, and generate a visual feature vector through hidden layer transformation, linear projection and pooling operations.

[0020] Optionally, the visual language model introduces a contrast loss function based on image-text pairing in the pre-training stage to enhance the alignment strength of image-text representations, introduces learnable cue word embedding parameters in the fine-tuning stage, and adds a fusion gating mechanism after the multi-head cross-attention layer to control the propagation ratio of image features in cross-modal generation.

[0021] Optionally, the S2 specifically includes:

[0022] S21, build a visual language model based on Qwen-VL architecture, perform data separation on the preprocessed image data, and obtain standardized image I s and a normalized text sequence T;

[0023] S22, normalize the image I s Input image encoding part, extract image features under multi-scale receptive field, generate image feature vector V through residual connection, layer normalization and linear mapping;

[0024] S23, input the text encoding part of the standardized text sequence T, and embed the position information using the sine and cosine position encoding function, and combine the self-attention mechanism to obtain the text semantics to generate the text embedding representation E;

[0025] S24. Input the image feature vector V and the text embedding representation E into the multi-head cross attention layer. After bidirectional alignment, the fused visual feature vector F is generated.

[0026] Optionally, the S3 specifically includes:

[0027] S31, construct prompt word template P = {p1, p2, ..., p l}, where p i is the i-th token in the prompt word, l is the length of the prompt word, and i∈{1,2,...,l};

[0028] S32, fine-tuning the prompt word, and inputting the prompt word and the visual feature vector F into the visual language model, and obtaining an intermediate fusion representation through a cross-attention mechanism;

[0029] S33, map the intermediate fusion representation through the fully connected layer and the bias term to obtain the structured text description sequence T d , the structured text description includes numerical information, data relationships and text elements.

[0030] Optionally, the prompt word template is constructed based on the corpus of a specific field, adopts a learnable embedding matrix and a positional encoding joint representation, performs context modeling on the initial prompt word sequence through a bidirectional GRU network, outputs the sequence as the initial prompt word embedding, and further dynamically fuses it with the image features through the attention weight to form an adaptive prompt word template with semantic enhancement capabilities, and jointly minimizes the image-text pairing loss function and the downstream task objective function during the training process to optimize the prompt word generation quality.

[0031] Optionally, the S4 specifically includes:

[0032] S41. Build a large language model based on the Transformer architecture, convert the structured text description into a vector matrix X, and directly concatenate the position encoding matrix P d , get the intermediate matrix X p =X+P d ;

[0033] S42, the intermediate matrix X p Input a multi-layer Transformer encoder and define the hidden state of each layer. The calculation form of the lth layer is:

[0034] H (l) =LN(H (l-1) +MHAtt(H (l-1) ))+FFN(H (l-1) );

[0035] Among them, H (l) represents the hidden state of the lth layer, l represents the layer index, LN(·) represents layer normalization, MHAtt(·) represents the multi-head attention mechanism, FFN(·) represents the feedforward neural network, and the initial state is H (0) =X p ;

[0036] S43. Introducing the LoRA low-rank decomposition method and two trainable low-rank matrices A and B to supervise fine-tune the weight parameters of the large language model:

[0037]

[0038] Where W represents the initial weight matrix and α represents the scaling factor, which is used to control the amplitude of the low-rank weight update.

[0039] S44, the intermediate matrix X pInput a large language model to generate a predicted probability distribution at each time step, and sample or decode it to obtain a candidate sequence Q for data analysis questions.

[0040] Optionally, the S5 specifically includes:

[0041] S51, the data analysis question candidate sequence Q and the structured text description sequence T d Encode and get the embedding matrix E Q With E T ;

[0042] S52, construct the attention matrix A, perform semantic consistency detection, and the element a in A ij It is defined as the attention weight between the i-th word in the candidate sequence and the j-th word in the text description:

[0043]

[0044] in, represents the row vector of the i-th word embedding in the candidate sequence, Represents the j-th word embedding row vector of the text description, Represents the j′th word embedding row vector of the text description, W a ,W b is the learnable projection matrix, h is the attention projection dimension, and exp(·) represents the natural exponential function with e as the base;

[0045] S53, generate context-aware embedding representation C according to the attention matrix A, and each row Where k is the number of tokens in the structured text description;

[0046] S54. Calculate the context aggregation vector v of the candidate sequence of data analysis questions based on the context-aware embedding representation C. q and structured text description embedding vector v t :

[0047]

[0048] Where n is the number of words in the candidate sequence;

[0049] S55, based on v q and v t , define the semantic consistency scoring function

[0050]

[0051] Among them, M is the learnable transformation matrix, ‖·‖ represents the two-norm of the vector, Indicates the degree of contextual semantic alignment;

[0052] S56. Set the semantic consistency threshold τ∈[0,1], if Then the candidate sequence Q of the data analysis question is judged as a semantically inconsistent sequence;

[0053] S57、For all satisfaction The sequence is used for structural rule matching detection, and the regular structure rule set is Match each rule in turn, and The sequences with errors are eliminated, and a set of data analysis question sequences that have passed semantic consistency detection and structural rule matching verification are output.

[0054] The beneficial effects of the present invention are:

[0055] First, the present invention constructs a visual language model based on the Qwen-VL architecture, adopts a dual-stream input structure of images and text and a multi-head cross-attention mechanism, and realizes efficient extraction and in-depth understanding of various types of visual information contained in the image, such as tables, graphics, numerical values, coordinate axis labels and text descriptions. It overcomes the technical bottleneck of the existing model's weak ability to understand complex data images and incomplete information extraction, and effectively improves the accuracy and stability of image semantic modeling.

[0056] Secondly, the invention introduces a dynamic construction method for structured prompt word templates. Through bidirectional GRU modeling and image feature fusion, it achieves a high degree of coupling between the semantic expression of prompt words and visual information. Combined with learnable embedding and gating mechanisms, this significantly enhances the adaptability of the visual language model for professional tasks and the quality of content generation. Furthermore, through supervised fine-tuning of a large language model and training with real-world data analysis corpus, the generated questions possess strong logical integrity, standardized terminology, and natural question expression, overcoming the lack of generalization of existing generation methods in professional scenarios.

[0057] Finally, the present invention strictly screens the generated question sequences by constructing a semantic consistency detection and structural rule matching mechanism to ensure the comprehensive consistency of the output questions in terms of language expression, structural form, and logical reasoning. At the same time, the use of image-text alignment triple encoding storage method ensures that the original image, structured description, and question output have strong correlation and traceability, providing a data foundation for subsequent multimodal task training and scenario deployment. The overall method improves the automation level and professionalism of data analysis question generation, and has good practical application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0059] Figure 1 This is a flow chart of the method for generating data analysis questions based on the combination of image input and large models proposed by the present invention;

[0060] Figure 2 This is a framework diagram of the large model combination architecture of the data analysis problem generation method based on image input and large model combination proposed by the present invention;

[0061] Figure 3 This is a schematic diagram of feature extraction and fusion of the data analysis problem generation method based on the combination of image input and large model proposed by the present invention;

[0062] Figure 4 A schematic diagram of generating a structured text description of a data analysis problem generation method based on a combination of image input and a large model proposed in the present invention; DETAILED DESCRIPTION

[0063] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0064] refer to Figure 1-4 ,The data analysis problem generation method based on the combination of image input and large model includes the following steps:

[0065] S1. Obtain original image data and perform preprocessing;

[0066] S2. Build a visual language model based on the Qwen-VL architecture, input the preprocessed image data into the visual encoder, extract image features and text semantics, and perform bidirectional alignment to generate a visual feature vector;

[0067] S3. Construct a prompt word template, input the prompt word into the visual language model through fine-tuning, and fuse it with the visual feature vector through a cross-attention mechanism to generate a structured text description containing numerical information, data relationships and text elements;

[0068] S4. Build a large language model based on the Transformer architecture and use the LoRA method for supervised fine-tuning. Take structured text descriptions as input sequences and generate candidate sequences of data analysis questions.

[0069] S5. Perform semantic consistency check and structural rule matching on candidate sequences of data analysis questions, eliminate sequences that do not meet semantic specifications or structural constraints, and output the data analysis question sequences that pass the check;

[0070] S6. Combine the original image data, structured text description, and data analysis question sequence into an image-text alignment triplet, and write it into a JSON format data structure for encoding and storage.

[0071] By constructing an end-to-end process combining image input and large models, the present invention realizes a complete closed loop from image acquisition, preprocessing, visual feature extraction, text fusion, question generation to structured storage, effectively improving the intelligence level and process efficiency of automatic generation of data analysis questions, and is suitable for multimodal understanding and task instruction-driven scenarios.

[0072] In this embodiment, the original image data is composed of visual data content, including tables, graphics, statistical figures, coordinate axis labels and text descriptions.

[0073] The present invention limits the image data source to visual statistical content, such as charts, numerical values ​​and text information, so that the system has clear application boundaries and data feature adaptation capabilities at the input layer, thereby improving the model's accuracy in understanding structural images and the targetedness of its generation.

[0074] In this embodiment, the preprocessing includes size normalization, color channel standardization and format conversion.

[0075] This paper introduces standard image preprocessing methods, including size normalization, color channel standardization and format conversion, which improves the consistency and stability of visual model input and optimizes the generalization performance of the model under diverse image inputs from the bottom layer.

[0076] In this embodiment, the Qwen-VL architecture adopts a dual-stream input structure of images and texts. The visual encoder includes two parts: image encoding and text encoding. In the image encoding part, image features under different receptive fields are extracted, and image feature vectors are generated through residual connection, layer normalization and linear mapping. In the text encoding part, position encoding and self-attention mechanism are used to obtain text semantics, generate text embedding representation, and perform bidirectional alignment of the image feature vector and the text embedding representation in the multi-head cross attention layer, and generate a visual feature vector through hidden layer transformation, linear projection and pooling operations.

[0077] The present invention adopts a dual-stream input structure of images and texts in the visual language model architecture, combines the residual connection and cross-attention mechanism of image encoding and text encoding, realizes the deep fusion between visual and language modalities, and enhances the model's ability to model the semantic relationship between images and texts.

[0078] In this embodiment, the visual language model introduces a contrast loss function based on image-text pairing in the pre-training stage to enhance the alignment strength of image-text representation, introduces learnable prompt word embedding parameters in the fine-tuning stage, and adds a fusion gating mechanism after the multi-head cross-attention layer to control the propagation ratio of image features in cross-modal generation.

[0079] The present invention optimizes the image-text alignment effect in the pre-training stage through the contrast loss function of image-text pairing, and introduces a fusion gating mechanism to control the intensity of image information propagation, achieving information selectivity and fidelity in multimodal content fusion, and improving the accuracy and relevance of generated text.

[0080] In this embodiment, S2 specifically includes:

[0081] S21, build a visual language model based on Qwen-VL architecture, perform data separation on the preprocessed image data, and obtain standardized image I s and a normalized text sequence T;

[0082] S22, normalize the image I s Input image encoding part, extract image features under multi-scale receptive field, generate image feature vector V through residual connection, layer normalization and linear mapping;

[0083] S23, input the text encoding part of the standardized text sequence T, and embed the position information using the sine and cosine position encoding function, and combine the self-attention mechanism to obtain the text semantics to generate the text embedding representation E;

[0084] S24. Input the image feature vector V and the text embedding representation E into the multi-head cross attention layer. After bidirectional alignment, the fused visual feature vector F is generated.

[0085] By extracting and aligning image features and text features separately in a multi-layer network, the present invention ensures fine control and semantic integrity in each step from image encoding, text position encoding to cross-attention fusion, thereby improving the expression quality of intermediate visual features.

[0086] In this embodiment, S3 specifically includes:

[0087] S31, construct prompt word template P = {p1, p2, ..., p l}, where p i is the i-th token in the prompt word, l is the length of the prompt word, and i∈{1,2,...,l};

[0088] S32, fine-tuning the prompt word, and inputting the prompt word and the visual feature vector F into the visual language model, and obtaining an intermediate fusion representation through a cross-attention mechanism;

[0089] S33, map the intermediate fusion representation through the fully connected layer and the bias term to obtain the structured text description sequence T d , the structured text description includes numerical information, data relationships and text elements.

[0090] This paper proposes an input method for structured prompt word templates, and fuses prompt words with visual features through a cross-attention mechanism, so that the generated structured text description can be closely organized around the image content, enhancing the controllability and professional adaptability of the model output.

[0091] In this embodiment, the prompt word template is constructed based on the corpus of a specific field, and adopts a learnable embedding matrix and position encoding joint representation. The initial prompt word sequence is contextually modeled through a bidirectional GRU network, and the output sequence is used as the initial prompt word embedding. It is further dynamically fused with image features through attention weights to form an adaptive prompt word template with semantic enhancement capabilities. During the training process, the image-text pairing loss function and the downstream task objective function are jointly minimized to optimize the prompt word generation quality.

[0092] The present invention constructs a prompt word template based on specific corpus, and combines learnable embedding, bidirectional GRU modeling and attention fusion technology to form an adaptive semantic enhancement prompt mechanism, which improves the guiding effect of prompt words on generation tasks and the model's ability to perceive task objectives.

[0093] In this embodiment, the S4 specifically includes:

[0094] S41. Build a large language model based on the Transformer architecture, convert the structured text description into a vector matrix X, and directly concatenate the position encoding matrix P d , get the intermediate matrix X p =X+P d ;

[0095] S42, the intermediate matrix X p Input a multi-layer Transformer encoder and define the hidden state of each layer. The calculation form of the lth layer is:

[0096] H (l) =LN(H (l-1) +MHAtt(H (l-1) ))+FFN(H (l-1) );

[0097] Among them, H (l) represents the hidden state of the lth layer, l represents the layer index, LN(·) represents layer normalization, MHAtt(·) represents the multi-head attention mechanism, FFN(·) represents the feedforward neural network, and the initial state is H (0) =Xp ;

[0098] S43. Introducing the LoRA low-rank decomposition method and two trainable low-rank matrices A and B to supervise fine-tune the weight parameters of the large language model:

[0099]

[0100] Where W represents the initial weight matrix and α represents the scaling factor, which is used to control the amplitude of the low-rank weight update.

[0101] S44, the intermediate matrix X p Input a large language model to generate a predicted probability distribution at each time step, and sample or decode it to obtain a candidate sequence Q for data analysis questions.

[0102] This paper introduces LoRA low-rank decomposition technology in the training process of large language models, combines the Transformer structure with positional encoding optimization strategy, and realizes efficient fine-tuning of large models under low resources, thereby improving the performance stability and training efficiency of the model in data analysis problem generation tasks.

[0103] In this embodiment, the S5 specifically includes:

[0104] S51, the data analysis question candidate sequence Q and the structured text description sequence T d Encode and get the embedding matrix E Q With E T ;

[0105] S52, construct the attention matrix A, perform semantic consistency detection, and the element a in A ij It is defined as the attention weight between the i-th word in the candidate sequence and the j-th word in the text description:

[0106]

[0107] in, represents the row vector of the i-th word embedding in the candidate sequence, Represents the j-th word embedding row vector of the text description, Represents the j′th word embedding row vector of the text description, W a ,W b is the learnable projection matrix, h is the attention projection dimension, and exp(·) represents the natural exponential function with e as the base;

[0108] S53, generate context-aware embedding representation C according to the attention matrix A, and each row Where k is the number of tokens in the structured text description;

[0109] S54. Calculate the context aggregation vector v of the candidate sequence of data analysis questions based on the context-aware embedding representation C. q and structured text description embedding vector v t :

[0110]

[0111] Where n is the number of words in the candidate sequence;

[0112] S55, based on v q and v t , define the semantic consistency scoring function

[0113]

[0114] Among them, M is the learnable transformation matrix, ‖·‖ represents the two-norm of the vector, Indicates the degree of contextual semantic alignment;

[0115] S56. Set the semantic consistency threshold τ∈[0,1], if Then the candidate sequence Q of the data analysis question is judged as a semantically inconsistent sequence;

[0116] S57、For all satisfaction The sequence is used for structural rule matching detection, and the regular structure rule set is Match each rule in turn, and The sequences with errors are eliminated, and a set of data analysis question sequences that have passed semantic consistency detection and structural rule matching verification are output.

[0117] The present invention proposes a semantic consistency detection scheme that combines attention mechanism, context perception and regular rule matching to achieve semantic alignment accuracy control and structural compliance verification of question candidate sequences, significantly improving the availability, rationality and professional quality of the final generated questions.

[0118] Example 1:

[0119] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the automatic generation module of data analysis questions of a certain educational intelligent learning platform. The platform provides data analysis question training functions for users preparing for exams. The traditional question setting method relies on teachers to manually write questions, which takes a lot of time and energy from image interpretation, data information extraction to question writing. Especially when dealing with questions with more complex image content such as statistical graphs, tabular data, composite graphs, etc., the accuracy of manual interpretation is limited, the generation efficiency is low, and the question structure and semantic quality are uneven, which makes it difficult to meet the needs of large-scale, high-quality question bank construction. Therefore, the platform introduces the data analysis question generation method based on the combination of image input and large models proposed by the present invention to realize the automatic generation from original images to standardized questions, and solve the problems of high labor costs, difficult quality control, and low update frequency.

[0120] The technical process provided by the embodiments of the present invention primarily includes the following steps. First, authoritative data image resources are collected online. This embodiment uses statistical chart samples from the "Data Release" section of the National Bureau of Statistics website as input. These chart types include bar charts, line charts, pie charts, and area charts. These images contain a large number of realistically meaningful numbers and contrast structures. After collection, the images are input into the system. During the preprocessing phase, the image size, color channels, and data format are unified to prepare for subsequent model input.

[0121] Next, the image is passed to the visual language model through the API. In this example, the Qwen-VL architecture model is used. After receiving the image input, the model generates a structured text description of the image based on the set prompt words. The prompt words have been optimized by domain experts and have a clear task-oriented orientation. For example: "You are an expert in the field of data science. Please generate a text description based on the data in the figure. You need to describe the data in general sentences and add specific numbers and proportions as appropriate. At the same time, you must also express trends, extreme values ​​and other features in the data. Do not include causal analysis or summary statements of data features. The word count cannot be less than 300 words." After fine-tuning the prompt words, the model can more accurately output natural language text descriptions containing key information such as image data relationships, numerical trends, and comparative changes.

[0122] Subsequently, the system inputs the above structured text into a supervised fine-tuned large language model for question generation. In this embodiment, the large language model used is Qwen1.5-7b, and the LoRA method is used for low-rank parameter tuning, so that the model can be efficiently adapted to data in specific fields, thereby improving the logical rigor and structural standardization of the generated text. The fine-tuning training data is mainly selected from civil servant information analysis test questions in recent years, and an input-output dialogue format is constructed, with the question stem as input and the question as output, to ensure that the model has the test language style, knowledge point coverage and expression logic. During the training process, only the parameters of the LoRA insertion layer are updated, which greatly reduces the consumption of training resources, and the generation effect is converged in less than 10 rounds of training.

[0123] The final question sequence output by the system undergoes semantic consistency checking and structural rule matching to ensure that the output meets the logical and formal requirements of the data analysis question at both the semantic and structural levels. After this complete process, the system integrates the image, text description, and generated question into a triple, which is stored and returned via the interface in a JSON structure, facilitating platform content access and user-friendly front-end display.

[0124] To verify the effectiveness of this method, we conducted a horizontal comparison with the existing mainstream GPT-4 API automatic generation method and manual template generation methods. Using 200 identical image samples, we generated corresponding data analysis questions and examined key metrics such as the generated BLEU score, ROUGE-L score, structural compliance, semantic consistency, and generation speed.

[0125] Table 1 Comparison of the effects of the automatic generation system of data analysis questions of the present invention and the existing methods

[0126]

[0127]

[0128] The results show that the solution of the present invention achieved a BLEU score of 44.8, which is higher than 38.3 of GPT-4 and 34.5 of the manual template; it reached 57.9 on ROUGE-L, which is significantly better than 51.2 of GPT-4 and 47.1 of the manual template; in terms of structural compliance, it reached 98.3%, close to 100% of the manual template and far exceeding 86.4% of GPT-4; in terms of semantic consistency, the present invention achieved an accurate matching rate of 94.7%, while GPT-4 was 83.8% and the manual template was 96.2%; in terms of average generation time, the present invention took 6.5 seconds, GPT-4 took 12.8 seconds, and the manual template took as long as 147 seconds due to manual participation, with a significant efficiency gap.

[0129] Finally, in a manual usability evaluation, 91% of the questions generated by this invention were rated as ready for immediate use, compared to 82% for GPT-4 and 85% for manual templates. Among the error samples, the logical error rate for questions generated by this invention was only 1.1%, compared to 4.4% for GPT-4 and 0.7% for manual templates. This demonstrates that while maintaining high-quality output, this invention also possesses extremely high generation stability and practicality, capable of replacing a significant amount of repetitive manual work in actual teaching and examination systems, improving the frequency of question updates and quality assurance.

[0130] In summary, this embodiment fully demonstrates that the solution of the present invention is highly practical and superior in the task of automatically generating data analysis questions. It not only improves the generation efficiency, but also comprehensively surpasses mainstream large models and traditional manual methods in terms of text quality, logical expression, and structural integrity, showing a broad prospect for promotion and application.

[0131] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A data analysis problem generation method based on a combination of image input and a large model, characterized in that: The steps include: S1. Obtain original image data and perform preprocessing; S2. Build a visual language model based on the Qwen-VL architecture, input the preprocessed image data into the visual encoder, extract image features and text semantics, and perform bidirectional alignment to generate a visual feature vector; S3. Construct a prompt word template, input the prompt word into the visual language model through fine-tuning, and fuse it with the visual feature vector through a cross-attention mechanism to generate a structured text description containing numerical information, data relationships and text elements; S4. Build a large language model based on the Transformer architecture and use the LoRA method for supervised fine-tuning. Take structured text descriptions as input sequences and generate candidate sequences of data analysis questions. S5. Perform semantic consistency check and structural rule matching on candidate sequences of data analysis questions, eliminate sequences that do not meet semantic specifications or structural constraints, and output the data analysis question sequences that pass the check; S6. Combine the original image data, structured text description, and data analysis question sequence into an image-text alignment triplet, and write it into a JSON format data structure for encoding and storage.

2. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The original image data consists of visual data content, including tables, graphics, statistical figures, coordinate axis labels and text descriptions.

3. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The preprocessing includes size normalization, color channel standardization and format conversion.

4. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The Qwen-VL architecture adopts a dual-stream input structure of image and text. The visual encoder includes two parts: image encoding and text encoding. In the image encoding part, image features under different receptive fields are extracted, and image feature vectors are generated through residual connection, layer normalization and linear mapping. In the text encoding part, position encoding and self-attention mechanism are used to obtain text semantics and generate text embedding representation. The image feature vector and the text embedding representation are bidirectionally aligned in the multi-head cross-attention layer, and the visual feature vector is generated through hidden layer transformation, linear projection and pooling operations.

5. The method for generating data analysis questions based on the combination of image input and large model according to claim 1, characterized in that: The visual language model introduces a contrast loss function based on image-text pairing in the pre-training stage to enhance the alignment strength of image-text representations. In the fine-tuning stage, it introduces learnable cue word embedding parameters and adds a fusion gating mechanism after the multi-head cross-attention layer to control the propagation ratio of image features in cross-modal generation.

6. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The S2 specifically includes: S21, build a visual language model based on Qwen-VL architecture, perform data separation on the preprocessed image data, and obtain standardized image I s and a normalized text sequence T; S22, normalize the image I s Input image encoding part, extract image features under multi-scale receptive field, generate image feature vector V through residual connection, layer normalization and linear mapping; S23, input the text encoding part of the standardized text sequence T, and embed the position information using the sine and cosine position encoding function, and combine the self-attention mechanism to obtain the text semantics to generate the text embedding representation E; S24. Input the image feature vector V and the text embedding representation E into the multi-head cross attention layer. After bidirectional alignment, the fused visual feature vector F is generated.

7. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The S3 specifically includes: S31, construct prompt word template P = {p1, p2, ..., p l }, where p i is the i-th token in the prompt word, l is the length of the prompt word, and i∈{1,2,...,l}; S32, fine-tuning the prompt word, and inputting the prompt word and the visual feature vector F into the visual language model, and obtaining an intermediate fusion representation through a cross-attention mechanism; S33, map the intermediate fusion representation through the fully connected layer and the bias term to obtain the structured text description sequence T d , the structured text description includes numerical information, data relationships and text elements.

8. The method for generating data analysis questions based on the combination of image input and large models according to claim 7, characterized in that: The prompt word template is constructed based on the corpus of a specific field, adopts a learnable embedding matrix and a positional encoding joint representation, performs context modeling on the initial prompt word sequence through a bidirectional GRU network, and outputs the sequence as the initial prompt word embedding. It is further dynamically fused with image features through attention weights to form an adaptive prompt word template with semantic enhancement capabilities. During the training process, the image-text pairing loss function and the downstream task objective function are jointly minimized to optimize the prompt word generation quality.

9. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The S4 specifically includes: S41. Build a large language model based on the Transformer architecture, convert the structured text description into a vector matrix X, and directly concatenate the position encoding matrix P d , get the intermediate matrix X p =X+P d ; S42, the intermediate matrix X p Input a multi-layer Transformer encoder and define the hidden state of each layer. The calculation form of the lth layer is: H (l) =LN(H (l-1) +MHAtt(H (l-1) ))+FFN(H (l-1) ); Among them, H (l) represents the hidden state of the lth layer, l represents the layer index, LN(·) represents layer normalization, MHAtt(·) represents the multi-head attention mechanism, FFN(·) represents the feedforward neural network, and the initial state is H (0) =X p ; S43. Introducing the LoRA low-rank decomposition method and two trainable low-rank matrices A and B to supervise fine-tune the weight parameters of the large language model: Where W represents the initial weight matrix and α represents the scaling factor, which is used to control the amplitude of the low-rank weight update. S44, the intermediate matrix X p Input a large language model to generate a predicted probability distribution at each time step, and sample or decode it to obtain a candidate sequence Q for data analysis questions.

10. The method for generating data analysis questions based on the combination of image input and large models according to claim 1, characterized in that: The S5 specifically includes: S51, the data analysis question candidate sequence Q and the structured text description sequence T d Encode and get the embedding matrix E Q With E T ; S52, construct the attention matrix A, perform semantic consistency detection, and the element a in A ij It is defined as the attention weight between the i-th word in the candidate sequence and the j-th word in the text description: in, represents the row vector of the i-th word embedding in the candidate sequence, Represents the j-th word embedding row vector of the text description, Indicates the text description j ′ word embedding row vector, W a ,W b is the learnable projection matrix, h is the attention projection dimension, and exp(·) represents the natural exponential function with e as the base; S53, generate context-aware embedding representation C according to the attention matrix A, and each row Where k is the number of tokens in the structured text description; S54. Calculate the context aggregation vector v of the candidate sequence of data analysis questions based on the context-aware embedding representation C. q and structured text description embedding vector v t : Where n is the number of words in the candidate sequence; S55, based on v q and v t , define the semantic consistency scoring function Among them, M is the learnable transformation matrix, ‖·‖ represents the two-norm of the vector, Indicates the degree of contextual semantic alignment; S56. Set the semantic consistency threshold τ∈[0,1], if Then the candidate sequence Q of the data analysis question is judged as a semantically inconsistent sequence; S57、For all satisfaction The sequence is used for structural rule matching detection, and the regular structure rule set is Match each rule in turn, and The sequences with errors are eliminated, and a set of data analysis question sequences that have passed semantic consistency detection and structural rule matching verification are output.

Citation Information

Cited By

  • Multi-modal data preprocessing and fusion technology and system based on artificial intelligence

    CN120833614A

  • Bidding document information extraction method

    CN121031593A

  • A method for extracting information from bidding documents

    CN121031593B

  • Geological disaster intelligent analysis method and system based on large-scale language model

    CN121456331A

  • Intelligent Analysis Method and System for Geological Hazards Based on Large-Scale Language Models

    CN121456331B