Image-text report generation method and device and storage medium

By classifying and preprocessing the signal analysis data, and using image and text processing models to generate graphic and text reports, the problems of inefficiency and poor quality in the existing technology are solved, and efficient and intelligent graphic and text report generation are achieved.

CN120448536AActive Publication Date: 2025-08-08LINGYANGE SEMICONDUCTOR, INC

Patent Information

Application Number
CN202510941425.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-08-08
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

The existing graphic and text report generation methods for signal analysis data are inefficient and poor in quality, and are easily disturbed by human factors.

Method used

By classifying and preprocessing the signal analysis data, feature vectors are extracted using image processing models and text processing models, and combined with search enhancement generation models, the image and text report is generated by fusion of images and text.

Benefits of technology

It realizes the precise fusion of images and text, and generates efficient and intelligent graphic reports with detailed contextual interpretation and logical association, improving the generation efficiency and professionalism.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448536A_ABST
    Figure CN120448536A_ABST
Patent Text Reader

Abstract

The invention provides an image-text report generation method and device and a storage medium, and the method comprises the steps: classifying signal analysis data according to a file type, carrying out the preprocessing, and determining the image data and text data of the signal analysis data; determining a first image processing model according to the image data, processing the image data by using the first image processing model, and determining a first image feature vector; determining a first text processing model according to the text data, and processing the first image feature vector and the text data by adopting the first text processing model to generate first text description; and fusing the first image feature vector and the first text description by adopting a retrieval enhancement generation model to generate a first image-text report of the signal analysis data. According to the image-text report generation method and the image-text report generation system, the generated image-text report can provide detailed context interpretation and logic association by adopting a retrieval enhancement generation model and a multi-model cooperative processing technology, so that the generation efficiency and the specialty of the image-text report are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and in particular to a method, device and storage medium for generating a graphic report. Background Art

[0002] In the field of data analysis technology, especially in multimodal data processing systems that combine images and text, an increasing number of applications require efficient and intelligent image and text generation. Traditional methods for generating graphic and text reports for signal analysis data typically rely on manual text input and require manual interpretation of the image data. This method is not only inefficient but also susceptible to human interference, resulting in poor quality graphic and text reports. Summary of the Invention

[0003] The present invention provides a method, device and storage medium for generating a graphic report, which are used to solve the problems of low efficiency and poor quality in the existing graphic report generation of signal analysis data.

[0004] Based on the above problems, in a first aspect, the present invention provides a method for generating a graphic report, comprising: Classifying the signal analysis data according to the file type and performing preprocessing to determine the image data and text data of the signal analysis data; Determining a first image processing model based on the image data, and processing the image data using the first image processing model to determine a first image feature vector; Determining a first text processing model based on the text data, and processing the first image feature vector and the text data using the first text processing model to generate a first text description; A retrieval-enhanced generation model is used to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data.

[0005] In conjunction with the first aspect, in a possible implementation, determining the first image processing model according to the image data includes: classifying the image data according to image metadata of the image data or using a lightweight classification model to determine the image data type; A first image processing model is determined according to the image data type and a matching relationship between the image data type and the image processing model.

[0006] In combination with the first aspect, in a possible implementation, the method further includes: determining the first image processing model in the following manner: Read the performance headroom of the computing system; A first image processing model is determined according to the performance margin of the computing system.

[0007] In conjunction with the first aspect, in a possible implementation manner, the first image feature vector includes: target detection region attributes and anomaly detection region attributes; The target detection area attributes include: target detection area type, target detection area location, target detection area size and target detection area label; The anomaly detection area attributes include: anomaly detection area type, anomaly detection area location, anomaly detection area size and anomaly detection area label.

[0008] In conjunction with the first aspect, in one possible implementation, determining the first text processing model according to the text data includes: Determining text processing model selection factors based on the text data; wherein the text processing model selection factors include: text type, text complexity and target output style; A first text processing model is determined according to the text processing model selection elements.

[0009] In conjunction with the first aspect, in one possible implementation, the retrieval-enhanced generation model is used to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data, including: Build a search knowledge base; Using the first image feature vector as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a first retrieval result; The first search result and the first text description are input into a first text processing model to generate a first graphic and text report of the signal analysis data.

[0010] In combination with the first aspect, in a possible implementation, the method further includes: determining a top-k value in the first search result according to the image data type.

[0011] In combination with the first aspect, in a possible implementation manner, the method further includes: Adjusting the correlation between the image and the text in the first graphic report to generate a second graphic report of the signal analysis data includes: Using an image understanding model to perform semantic back-matching on the first image-text report to determine a consistency assessment result; If the consistency assessment result is lower than a preset score, a semantic reasoning model is used to determine the missing text content in the first graphic report; Using the missing content of the text as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a second retrieval result; The second search result and the first graphic report are input into a first text processing model to generate a second graphic report of the signal analysis data.

[0012] In a second aspect, a graphic report generating device is provided, comprising: A preprocessing module, configured to classify the signal analysis data according to file type, and perform preprocessing to determine image data and text data of the signal analysis data; an image processing module, configured to determine a first image processing model based on the image data, and process the image data using the first image processing model to determine a first image feature vector; a text processing module, configured to determine a first text processing model based on the text data, and process the first image feature vector and the text data using the first text processing model to generate a first text description; A fusion module is used to adopt a retrieval enhancement generation model to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data.

[0013] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the graphic report generation method as described in the first aspect or any possible implementation method in combination with the first aspect are executed.

[0014] The beneficial effects of the present invention include: The method, device and storage medium for generating a graphic report provided by the present invention include: classifying signal analysis data according to file type and performing preprocessing to determine image data and text data of the signal analysis data; determining a first image processing model according to the image data, and using the first image processing model to process the image data to determine a first image feature vector; determining a first text processing model according to the text data, and using the first text processing model to process the first image feature vector and text data to generate a first text description; using a retrieval-enhanced generation model to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data. The method for generating a graphic report provided by the present invention, by combining the first image processing model with the first text processing model, and using the retrieval-enhanced generation model and multi-model collaborative processing technology, can more accurately process the association between image data and text data, deepen contextual understanding, and the generated graphic report can provide detailed contextual explanations and logical associations, so that images and texts are accurately integrated, thereby improving the generation efficiency and professionalism of the graphic report and realizing an efficient and intelligent graphic generation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 A flowchart of the method for generating a graphic report provided by the present invention; Figure 2 This is a structural diagram of the graphic report generation device provided by the present invention. DETAILED DESCRIPTION

[0016] The present invention provides a method, apparatus, and storage medium for generating a graphic report. Preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are intended only to illustrate and explain the present invention and are not intended to limit the present invention. Furthermore, the embodiments and features of the embodiments herein may be combined unless there is a conflict.

[0017] The present invention provides a method for generating a graphic report, such as Figure 1 Shown, including: S101, classifying the signal analysis data according to the file type and performing preprocessing to determine the image data and text data of the signal analysis data; S102, determining a first image processing model based on the image data, and processing the image data using the first image processing model to determine a first image feature vector; S103, determining a first text processing model based on the text data, and using the first text processing model to process the first image feature vector and the text data to generate a first text description; S104: Using a retrieval-enhanced generation model, the first image feature vector and the first text description are merged to generate a first graphic report of the signal analysis data.

[0018] The present invention, in the field of data analysis technology, is particularly aimed at scenarios of multimodal data processing systems that combine images and text. Signal analysis data can refer to the feature extraction, analysis and interpretation of various signals to reveal the information contained in various signals. Traditional methods for generating graphic reports of signal analysis data usually rely on manual text input to interpret image data. This method not only has the problem of low efficiency, but is also easily interfered by human factors, and the quality of the generated graphic reports is poor. With the evolution of artificial intelligence technology, especially the development of image recognition, natural language processing (NLP) and multimodal learning, more and more systems are trying to combine images and text. However, existing artificial intelligence technology still faces many technical challenges in the field of graphic report generation for signal analysis data, especially in terms of accurate fusion of image and text content, context understanding and processing efficiency.

[0019] In the present invention, signal analysis data may include images and text. For example, images may include spectrum diagrams, waveform diagrams, grayscale diagrams, color diagrams, heat maps, etc.; texts may include analysis reports, technical specifications, etc. According to the file type of the signal analysis data, such as the file extension, images and texts are classified. For example, .jpg and .png represent images, and .txt and .docx represent texts. The image is preprocessed, for example, through denoising, normalization, and resizing operations, so that the image can adapt to the input of the subsequent first image processing model, thereby obtaining image data of the signal analysis data. The text is preprocessed, for example, by performing operations such as word segmentation, stop word removal, and stem extraction on the text to ensure that the text content can be accurately understood by the first text processing model to obtain text data. The image data can be represented as: D image ={I1, I2, ..., I n}, where I1, I2, …, I n Represents each preprocessed image. Text data can be represented as: D text ={T1, T2, ..., T n}, where T1, T2, …, T nRepresents each preprocessed text. Based on the image data and text data, the first image processing model is used to process the image data. The first image processing model can be selected according to the type of image data. For example, if the image data belongs to medical imaging, the U-Net, SegFormer, and SAM models can be used to segment the image, and the YOLOv8 model can be used for target detection. The type and size of the target detection area are extracted from the image data, and the label of the target detection area is generated. The formula for determining the first image processing model is described as: M image =YOLO or Segformer. Thus, the first image feature vector is obtained. The formula of the first image feature vector is expressed as: F image =M image (D image ), where M image Represents the first image processing model. The first text processing model is used to process the first image feature vector and text data to generate a first text description. The first text processing model can be determined based on the characteristics of the text data and the model. For example, the biomedical generative pre-trained Transformer model (BioGPT, Generative Pre-trained Transformer for Biomedical Text Generation and Mining) is applicable to the medical field. When a large number of medical field terms are reflected in the text data, the BioGPT model can be selected as the first text processing model. The formula for determining the first text processing model is described as: M text =LLaMA or GPT; where LLaMA (Large Language Model Application) and GPT (Generative Pre-trained Transformer) represent optional first text processing models. The first text processing model is used to generate a first text description related to the image content based on the first image feature vector and text data. The first text description may include: a text description of the image content, a text description combining the image content and text data. The formula for the first text description is: R text =M text (F image , D text ), where M text Represents the first text processing model.

[0020] Furthermore, the Retrieval-Augmented Generation (RAG) model is used to deeply fuse the first image feature vector with the first text description to ensure high consistency and contextual relevance between the image and text content. The RAG module uses the first image feature vector as retrieval information to ensure that the generated text description is closely related to the details of the image content, guiding the first text processing model to generate an accurate first image and text report. The formula for generating the first image and text report is expressed as: R fusion =RAG(F image , R text ), where RAG represents the deep fusion of the first image feature vector and the first text description using a retrieval-augmented generative model. The first graphic and text report can be output in multiple formats, such as HTML and PDF, and can be viewed and downloaded.

[0021] By combining the first image processing model with the first text processing model, and adopting the retrieval-enhanced generation model and multi-model collaborative processing technology, the present invention can more accurately process the association between image data and text data, deepen contextual understanding, and the generated graphic report can provide detailed contextual explanations and logical associations, so that the image and text are accurately integrated, thereby improving the generation efficiency and professionalism of the graphic report, and realizing an efficient and intelligent graphic generation method.

[0022] In yet another embodiment of the present invention, determining the first image processing model based on the image data includes: Step 1: classify the image data according to the image metadata of the image data or use a lightweight classification model to determine the image data type; Step 2: Determine the first image processing model based on the image data type and the matching relationship between the image data type and the image processing model.

[0023] The present invention first classifies image data based on image metadata or a lightweight classification model to determine the image data type. Based on the pre-set matching relationship between the image data type and the image processing model, the most appropriate first image processing model is selected to achieve precise matching between the model and the image data. Regarding step 1 above, image metadata can refer to descriptive information inherent in the image data. Image metadata can include: image file name, path tag, and Exchangeable Image File Format (EXIF) information. The image file name includes the image naming information. The path tag includes the image storage path. EXIF information includes camera parameters, capture time information, capture location information, and content tags. Camera parameters may include resolution, color depth, format (such as JPG / PNG), slice thickness for medical CT scans, and exposure time for industrial cameras. Content tags may include category identifiers generated through manual annotation or file naming rules, such as "lung CT" and "gear defect image." The image data type can be determined by parsing key fields in the image metadata. For example, if the image metadata includes the "DICOM" format identifier and the content tag is "brain," the image data type is determined to be medical imaging data. A neural network with a small number of parameters and low computational complexity (such as MobileNet and ShuffleNet) is used. After lightweight training, it focuses on the coarse-grained classification of image types to obtain a lightweight classification model. A lightweight classification model is used to classify image data. The lightweight classification model can extract basic image distribution features such as contours, textures, CT grayscale, X-ray radiation patterns, industrial edge intensity, etc. through shallow convolutional layers, and output classification probabilities. For example, the probability of 60% is "industrial inspection image" and 30% is "educational image". The category with the highest probability is selected as the image data type. The image data type can represent the field in which the image is located, such as medical imaging, industrial inspection images, educational images, and general visual images. For the above step 2, a mapping table of "matching relationship between image data type and image processing model" is established. For example, it can be generated based on historical experimental data or industry experience. For example, it is shown in Table 1:

[0024] Table 1 The first image processing model is determined based on the image data type and the matching relationship between the image data type and the image processing model. If multiple corresponding image processing models are available, additional filtering parameters can be added, such as prioritizing YOLOv8 for processing speed and Faster R-CNN for accuracy, to determine the first image processing model. The first image processing model is customized based on the image data type. For example, medical imaging and industrial inspection require different models due to feature differences. Adapting the model can improve feature extraction accuracy. Key features of different image types (such as lesion texture in medical images and defect outlines in industrial images) require specific model structures (such as convolution kernel size and attention mechanism in CNN). The first image processing model can enhance feature extraction capabilities in a targeted manner. This eliminates the need for manual debugging or traversal of all models. Automatic determination through classification and matching rules reduces reliance on manual experience and trial-and-error costs, making it suitable for standardized image processing scenarios.

[0025] In another embodiment of the present invention, the first image processing model is determined in the following manner: Step 1: Read the performance margin of the computing system; Step 2: Determine a first image processing model based on the performance margin of the computing system.

[0026] The present invention first reads the computing system's performance headroom, such as CPU / GPU computing power and memory capacity. It then matches the corresponding image processing model based on specific metrics of the computing system's performance headroom, such as the percentage of remaining computing power and the amount of free memory, to achieve dynamic adaptation of model complexity to system resources. For step 1 above, the computing system's performance headroom can be accessed through the computing system API. The computing system's performance headroom includes CPU idle rate, GPU idle rate, and free physical memory. For step 2 above, a first image processing model is determined based on the computing system's performance headroom. For example, when real-time processing is required or when running on a resource-limited computing system, a lighter-weight image processing model (such as MobileNet) is used, which offers higher computational efficiency and is suitable for edge computing applications. Selecting the first image processing model based on the real-time computing system's performance headroom avoids resource waste caused by using a "large model for a small task" (e.g., using ResNet for simple image classification) or performance bottlenecks caused by using a "small model for a large task" (e.g., using MobileNet for high-resolution medical imaging). When resources are limited (e.g., performance headroom falls below a threshold), a lightweight model is automatically selected to prevent system overload and crashes. This is particularly useful for edge computing devices or multi-tasking scenarios, ensuring system stability. High-performance systems can use complex models to improve accuracy, while low-performance systems rely on lightweight models to ensure real-time performance. For example, when the GPU idle rate is >70%, the Transformer model is selected, while when it is <30%, a lightweight CNN variant is switched to improve processing efficiency. Furthermore, model deployment requires no manual intervention, and automatic adaptation to resource status reduces manual configuration workload and supports automated management of large-scale distributed computing environments.

[0027] In yet another embodiment of the present invention, the first image feature vector includes: target detection region attributes and anomaly detection region attributes; The target detection area attributes include: target detection area type, target detection area location, target detection area size and target detection area label; Anomaly detection area attributes include: anomaly detection area type, anomaly detection area location, anomaly detection area size, and anomaly detection area label.

[0028] In the present invention, the first image feature vector can be a structured description of key regions in the image data, including target detection region attributes and anomaly detection region attributes. Target detection region attributes are used to characterize the characteristics (type, location, size, and label) of normal targets in the image, while anomaly detection region attributes focus on the characteristic description of abnormal regions. Together, they constitute a semantic feature representation of the image, providing structured data support for subsequent analysis. The target detection region type can describe the physical category or functional attributes of the detection target. For example, "gear," "bolt," and "circuit board" in industrial inspection, and "lung," "liver," and "tumor" in medical imaging. The target detection region position represents the spatial coordinates of the target in the image, typically represented by a bounding box. For example, (x1, y1, x2, y2) represents the position coordinates of the upper left and lower right corners. The target detection region size represents the geometric dimensions of the target region. Target detection region labels can be standardized codes or text descriptions of target types, conforming to a predefined labeling system. For example, industrial inspection image labels might include {"gear":"GEAR_001","bolt":"BOLT_002"}; medical image labels might include {"lung":"LUNG_01"}. Target detection region labels ensure unified storage and retrieval of detection results from different images and image processing models. For example, batch retrieval of all gear images using the label "GEAR_001" is possible. Target detection region attributes represent objects that should be present, while anomaly detection region attributes represent defects that should not be present. For example, if the target detection region type in the target detection region attributes is "circuit board," the anomaly detection region type in the anomaly detection region attributes might be "defective solder joints on a circuit board." Anomaly detection region attributes can be applied to report generation. For example, in medical image report generation, anomaly detection region attributes can assist in generating diagnostic explanations. For example, they can be used to locate tumors, hemorrhages, lesions, and organ abnormalities. An example report generated might be, "A high-density shadow with irregular edges appears in the left lower lobe of the lung, suspected to be a nodular lesion." In industrial inspection report generation, anomaly detection region attributes can assist in generating defect detection reports. For example, they can be used to identify surface cracks, scratches, foreign matter, or structural defects. An example report generated might be, "A linear crack approximately 2.1mm long and 0.3mm deep appears in the upper left corner of the aluminum component." In educational image report generation, anomaly detection region attributes can be used to highlight key areas in textbooks (such as abnormal cell structure or structural errors) and automatically generate explanations or comments. In general vision report generation, anomaly detection region attributes can be used to identify intruders, illegal parking, and unusual fire sources. An example report generated might be, "An unauthorized person was detected in this area for more than 30 seconds."Abnormal detection areas may include: areas that differ from expectations, abnormal areas identified by model predictions, areas that highly overlap with manually annotated anomalies, and areas with low model prediction confidence. Areas that differ from expectations exhibit significant statistical differences from normal samples, such as deviations in grayscale, texture, or shape, asymmetric areas in brain images, and discontinuous copper traces in PCBs. For abnormal areas identified by model predictions, for example, classifiers and segmentation models are used to detect pixels or regions in the image data, marking areas with defects or abnormalities. The YOLOv8 model is used to annotate specific defects, such as cracks, with bounding boxes. For areas that highly overlap with manually annotated anomalies, comparison is performed with ground truth annotations, such as those from radiologists or quality control personnel, to identify areas with high overlap. For areas with low model prediction confidence, model uncertainty analysis methods, such as Monte Carlo Dropout (MC Dropout) or model ensembles, are used to identify areas with low model confidence. The anomaly detection region type describes the specific form or cause of the anomaly, such as "crack," "wear," "deformation," and "missing parts" in industrial inspections; and "nodule," "hemorrhage," and "fibrosis" in medical imaging. The anomaly detection region position follows the same format as the target detection region position. For example, in a gear image, the anomaly position is (200, 250, 220, 270), indicating a defect at the gear tooth tip. The anomaly detection region size represents the geometric dimensions of the anomaly region, such as crack length (e.g., 10mm), wear depth (e.g., 0.5mm), or the proportion of the anomaly to the target area (e.g., the area of a poor solder joint accounts for 20% of the total solder joint area). The anomaly detection region label represents a standardized code for the anomaly type, which may include severity or treatment recommendations. For example, the labels for industrial inspection images are: {"crack - mild": "CRACK_01", "crack - severe": "CRACK_02"}; and the labels for medical images are: {"benign nodule": "NODULE_BENIGN", "malignant nodule": "NODULE_MALIGNANT"}. The formula for determining the attributes of the anomaly detection area is expressed as: R. anomalous =M detect (F image ), where M detect Represents the detection module of the first image processing model. The first image processing model is used to process image data. The determined first image feature vector includes attributes of target detection regions and attributes of anomaly detection regions. This allows the model to handle both "target recognition" and "anomaly detection" tasks simultaneously. For example, in medical images, it can label both organ locations (target detection) and lesion regions (anomaly detection), enabling complex analysis.

[0029] In yet another embodiment of the present invention, determining a first text processing model based on text data includes: Step 1: Determine the text processing model selection factors based on the text data; the text processing model selection factors include: text type, text complexity, and target output style; Step 2: Select elements according to the text processing model to determine the first text processing model.

[0030] The present invention first extracts text processing model selection elements for text type, text complexity, and target output style from text data, and then matches the first text processing model to the corresponding elements based on the text processing model selection elements. Regarding step 1 above, text type can refer to the field and genre category to which the text content belongs. Examples include abstracts, teaching materials, and medical interpretations. This can be determined through text keyword recognition, topic model analysis, or predefined text classifiers. For example, if words such as "lung" and "tumor" frequently appear in a text, it can be determined to be a medical interpretation. Text complexity can measure the complexity of the text in terms of semantics, grammar, and logical structure, as well as the degree of association between images and text. Natural language processing techniques are used for analysis. For example, the complexity of the text is measured by counting the proportion of long sentences, the number of low-frequency words, the referential relationships between sentences, and the number of images associated with the text. Complexity is quantified by calculating a readability metric (such as the Flesch-Kincaid readability score). A lower score indicates a more difficult text to understand and a higher complexity. The target output style can refer to the language style and expression characteristics that the user expects the generated text to present, such as professional terminology, easy to understand, colloquial, educational, scientific, etc. The target output style can be explicitly specified through user input or inferred based on the purpose of the text and the target audience. Regarding step 2 above, the first text processing model is determined based on the text processing model selection factors. A correspondence table between "text processing model selection factors and first text processing models" is constructed, for example, as shown in Table 2:

[0031] Table 2 Based on experiments and practical application experience, the most suitable first text processing model for different combinations of text processing model selection factors can be determined. When there are multiple candidate first text processing models, further screening can be carried out based on other factors, such as the computing resource requirements, processing speed, and diversity of generated text of the first text processing model. For example, on a device with limited resources, lightweight models with a small number of parameters and fast running speed are preferred; if the diversity and innovation of the text are pursued, a model with stronger generation capabilities is selected. Selecting an appropriate first text processing model based on the text processing model selection factors can ensure that the model adapts to the subject and structural characteristics of the text. For example, when processing academic papers, choosing a model that excels at logical reasoning and professional terminology can improve the accuracy of text processing.

[0032] In another embodiment of the present invention, a retrieval-enhanced generation model is used to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data, including: Step 1: Build a search knowledge base; Step 2: Using the first image feature vector as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a first retrieval result; Step 3: Input the first search result and the first text description into a first text processing model to generate a first graphic report of the signal analysis data.

[0033] The present invention first constructs a retrieval knowledge base encompassing domain knowledge of signal analysis data. Then, using a first image feature vector as a retrieval condition, a retrieval-enhanced generative model is used to match relevant information from the retrieval knowledge base to obtain a first retrieval result. Finally, the first retrieval result and the first text description are input into a first text processing model to fuse and generate a structured first graphic and text report, achieving a deep integration of domain knowledge with image and text data. Regarding step 1 above, the retrieval knowledge base can integrate authoritative knowledge such as standards and historical cases in the field of signal analysis data to prevent the model from generating erroneous conclusions based solely on image feature vectors and text data. The retrieval knowledge base can include: typical image data anomaly samples, expert report fragments containing annotation information, standard sentence templates, a domain term ontology library, and multimodal comparison samples. Typical image data anomaly samples can refer to image samples in an image dataset that exhibit significant differences from normal images. These differences may be due to errors or interference during data acquisition, or to abnormal conditions inherent in the image itself. Expert report fragments containing annotation information can refer to report content that has been reviewed and annotated by experts. Annotation information typically includes classification, key points, and explanations of the report content. Standard sentence templates can refer to pre-designed sentence structure templates that are tailored to specific domains or purposes. These templates are typically used to generate standardized text content to ensure consistent and standardized information presentation. A domain term ontology can refer to a terminology system constructed for a specific domain, including specialized vocabulary, concepts, and their interrelationships. A domain term ontology not only lists terms but also defines the hierarchical structure and relationships between terms. Multimodal comparison samples can refer to data samples that include multiple modalities (such as text and images) and have corresponding relationships between these modalities. Multimodal comparison samples are used to study the association and consistency between data from different modalities. Regarding step 2 above, a pre-trained cross-modal model, such as the Contrastive Language-Image Pre-training model, is used to convert the first image feature vector into a semantic vector. This semantic vector is used as retrieval information, and a retrieval-enhanced generative model is used to search the retrieval knowledge base. A top-k search is performed based on semantic similarity to obtain the first retrieval result. Regarding step 3 above, the first retrieval result and the first text description are input into the first text processing model. The first text description includes feature interpretation based on image data and text data. The first search result combines the image data with background knowledge in the search knowledge base, thereby fusing the first image feature vector and the first text description. A first graphic and text report of the signal analysis data is generated through the first text processing model.For example, if the image data is a CT image, the abnormality detection region attribute of the first image feature vector is "a high-density area in the left lower lobe, 1.4 cm in diameter," and the first text description is "a high-density area with a diameter of 1.4 cm and blurred edges appears in the left lower lobe." Using the first image feature vector as retrieval information, the search knowledge base is searched using the retrieval-enhanced generative model. The first search results include "a shadow with a diameter of 1.2 cm appears in the left upper lobe; further CT scan is recommended" and "CT scan revealed ground-glass opacity in the right middle lobe, indicating an early-stage lesion." This first search result and the first text description are input into the first text processing model, generating the first graphic report of the signal analysis data, which reads, "Based on the CT image, a high-density shadow with a diameter of approximately 1.4 cm and blurred edges was observed in the left lower lobe. Further imaging is recommended to rule out early-stage lesions." The retrieval-enhanced generative model not only fuses the first image feature vector and the first text description, but also uses the retrieval knowledge base's enhancement mechanism to ensure that the generated text content is highly consistent with the image content. The retrieval-enhanced generation model guides the first text processing model through the first image feature vector, allowing the generated first graphic report to be dynamically adjusted based on the image content and context, thereby generating a more accurate description. Through retrieval enhancement of the first image feature vector, the contextual consistency between the image and text description is ensured. The generated text is no longer a simple description of the image content, but is intelligently generated based on the details and context in the image. Through intelligent reasoning and enhancement, the retrieval-enhanced generation model can dynamically adjust the generated text content based on the image content, making the first graphic report more in line with actual scenario needs and improving accuracy, logic, and professionalism.

[0034] In another embodiment of the present invention, the method further includes: determining a top-k value in the first search result according to the image data type.

[0035] The present invention determines the top-k value of the first search result (i.e., returns the top k search results) based on the image data type (e.g., medical images, industrial inspection images). The top-k value indicates that the search-enhanced generative model searches the search knowledge base and selects the k most relevant results from a large number of candidate data based on a similarity metric. Image data types may include medical images, industrial inspection images, educational images, and general visual images. Medical images can be generated using various imaging technologies for medical diagnosis, treatment, and research, and can reflect the internal structure and function of the human body. Industrial inspection images can be generated using various imaging technologies during industrial production to monitor product quality, equipment status, and production processes. Educational images can be images used for educational and teaching purposes, including textbook illustrations, teaching demonstrations, and experimental images. These images help students better understand and master knowledge. General visual images can be images commonly seen in daily life and not specific to a particular field. These images can be used for various visual tasks, such as image recognition, classification, and retrieval. Different image data types are adapted to different top-k values due to differences in domain knowledge complexity and feature diversity, achieving adaptive optimization of the number of search results. Setting the top-k value can optimize accuracy and coverage. For example, for medical images, a top-k value of 3 to 5 is recommended to improve accuracy. For industrial inspection images, educational images, and general visual images, a top-k value of 5 to 10 is recommended to improve coverage.

[0036] In another embodiment of the present invention, the method further comprises: Adjust the correlation between the image and text in the first graphic report to generate a second graphic report of the signal analysis data, including: Step 1: Use the image understanding model to perform semantic back-matching on the first image and text report to determine the consistency assessment result; Step 2: If the consistency assessment result is lower than the preset score, a semantic reasoning model is used to determine the missing text content in the first graphic report; Step 3: Using the missing content of the text as the search information, the search knowledge base is searched using the search enhancement generation model to determine the second search result; Step 4: Input the second search result and the first graphic report into the first text processing model to generate a second graphic report of the signal analysis data.

[0037] The present invention adopts a retrieval enhancement generation model, adjusts the correlation between images and texts through feedback mechanism and reasoning process, optimizes the first graphic report, and generates a second graphic report of signal analysis data. The formula for generating the second graphic report is expressed as: R enhanced =Inference(R fusion ); where R fusionRepresents the first image-text report; Inference represents the feedback mechanism and reasoning process. For the above step 1, the image understanding model is used to perform semantic back-matching on the first image-text report to determine whether the text description in the first image-text report can correctly correspond to the image content and determine the consistency assessment result. image , using the image understanding model to generate image embedding vectors , for the text description of the first graphic report R fusion , the extracted text embedding vector The formula for the consistency assessment result is described as:

[0038] in, Cosine similarity is used to measure the semantic consistency between the image and the text. For example, the CLIP and BLIP-2 models are used to perform semantic back-matching on the first image and text report. By calculating the similarity of the semantic features of the image and text, the consistency of the image and text is quantitatively scored. For example, if the signal frequency peaks displayed in the image closely match the frequency data mentioned in the text, the consistency score is high; conversely, if an abnormal area appears in the image but is not mentioned in the text, the score is low. The final output is the consistency assessment result, which reflects the close connection between the image and text content.

[0039] For step 2 above, if the consistency assessment result is lower than the preset score, the semantic matching of the first graphic report is insufficient or the content logic is broken. The semantic reasoning model is used to determine the missing text content in the first graphic report. For example, mask language modeling is performed on the text description of the first graphic report to determine whether the text description is missing words. In the first graphic report, if there is a missing part in the text, such as "The image shows increased lung density, [MASK]", the semantic reasoning model will automatically generate or complete it according to the image corresponding context and the parameters obtained by training. Predict the most likely missing words (such as "left lung", "nodule", "size", etc.). The formula for missing content in the text is described as:

[0040] in, The text is missing content; represents the vocabulary, i.e., the set of all possible words that the semantic reasoning model can predict. The context corresponding to the image in the given first graphic report and semantic reasoning model parameters Under the condition that Probability of occurrence. To find the probability The largest vocabulary .

[0041] Another example is detecting key regions in image data to identify areas not mentioned in the text description. Another example is using templates / knowledge to guide language model question answering, such as "Does this CT image mention the lesion on the left side?" to identify missing content in the text.

[0042] For step 3 above, the missing content of the text is used as the search information, and the search knowledge base is searched again using the search enhancement generation model to improve the professionalism of the missing content of the text and determine the second search result. The formula for the second search result is described as:

[0043] in, To search the knowledge base; To retrieve the vector representation of text in the knowledge base; is the vector representation of the missing content of the text.

[0044] For step 4 above, the second search results are integrated with the text content in the first graphic report, and inserted or supplemented to the appropriate position in a logical order to make the text content more complete and the logic more coherent. Ensure that the correspondence between the image and the newly added text content is clear. Input the integrated data into the first text processing model, and the model performs language optimization, logical sorting and formatting on the text, such as adjusting the fluency of sentences, adding transition sentences, and standardizing professional terminology. Finally, a second graphic report is generated that includes complete graphic information and is closely related, presenting the signal analysis data in a clear, accurate and professional form. The formula for the generated second graphic report is described as:

[0045] in, The context corresponding to the image in the second graphic report, is the first text processing model, is the vector representation of the second search result; is a parameter for the first text processing model. For example, the image data is a lung CT scan, and the description in the first graphic report reads, "Images show increased lung density; reexamination recommended." Using the image understanding model to perform semantic back-matching on the first graphic report, the image understanding model found no mention of "left lung" or "definite lesion size," resulting in a consistency assessment result below the preset score. The CLIP model indicates a nodular abnormality in the left middle lobe of the lung. The semantic reasoning model then asks, "Is this a nodule? What is its size?" to identify missing text content. Using the missing text content as retrieval information, the search-enhanced generative model searches the retrieval knowledge base to determine a second search result. The second search result and the first graphic report are then input into the first text processing model to generate a second graphic report of the signal analysis data. The second graphic report is optimized to produce, "A high-density nodule approximately 1.4 cm in diameter is present in the left middle lobe of the lung; further examination is recommended to exclude early-stage lesions." Through consistency assessment and missing content supplementation, the signal features presented in the image in the second graphic report are consistent with the text description, avoiding situations where abnormal waveforms are displayed in the image but not explained in the text, allowing readers to more clearly and accurately understand the signal analysis results. The coherent and complete second graphic report reduces the difficulty for readers to understand signal analysis data, reduces confusion caused by information gaps or contradictions, and improves the readability and practicality of the report.

[0046] Based on the same inventive concept, the present invention also provides a graphic report generation device. Since the principles of the problems solved by these devices are similar to those of the aforementioned graphic report generation method, the implementation of the device can refer to the implementation of the aforementioned method, and the repeated parts will not be repeated.

[0047] The present invention provides a graphic report generating device, such as Figure 2 Shown, including: A preprocessing module 201 is used to classify the signal analysis data according to the file type and perform preprocessing to determine the image data and text data of the signal analysis data; An image processing module 202 is configured to determine a first image processing model based on the image data, and process the image data using the first image processing model to determine a first image feature vector; a text processing module 203 for determining a first text processing model based on the text data, and processing the first image feature vector and the text data using the first text processing model to generate a first text description; The fusion module 204 is configured to fuse the first image feature vector and the first text description using a retrieval enhancement generation model to generate a first graphic and text report of the signal analysis data.

[0048] In another embodiment of the present invention, the image processing module 202 is configured to classify the image data according to image metadata of the image data or by using a lightweight classification model to determine the image data type; A first image processing model is determined according to the image data type and a matching relationship between the image data type and the image processing model.

[0049] In another embodiment of the present invention, the image processing module 202 is further configured to determine the first image processing model in the following manner: Read the performance headroom of the computing system; A first image processing model is determined according to the performance margin of the computing system.

[0050] In yet another embodiment of the present invention, the first image feature vector includes: target detection region attributes and anomaly detection region attributes; The target detection area attributes include: target detection area type, target detection area location, target detection area size and target detection area label; The anomaly detection area attributes include: anomaly detection area type, anomaly detection area location, anomaly detection area size and anomaly detection area label.

[0051] In another embodiment of the present invention, the text processing module 203 is configured to determine text processing model selection factors based on the text data; wherein the text processing model selection factors include: text type, text complexity, and target output style; A first text processing model is determined according to the text processing model selection elements.

[0052] In another embodiment of the present invention, the fusion module 204 is used to construct a search knowledge base; Using the first image feature vector as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a first retrieval result; The first search result and the first text description are input into a first text processing model to generate a first graphic and text report of the signal analysis data.

[0053] In another embodiment of the present invention, the fusion module 204 is further configured to determine a top-k value in the first search result according to the image data type.

[0054] In another embodiment of the present invention, the fusion module 204 is further configured to: Adjusting the correlation between the image and the text in the first graphic report to generate a second graphic report of the signal analysis data includes: Using an image understanding model to perform semantic back-matching on the first image-text report to determine a consistency assessment result; If the consistency assessment result is lower than a preset score, a semantic reasoning model is used to determine the missing text content in the first graphic report; Using the missing content of the text as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a second retrieval result; The second search result and the first graphic report are input into a first text processing model to generate a second graphic report of the signal analysis data.

[0055] Based on the same inventive concept, the present invention provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for generating a graphic report as described in any of the above embodiments are executed.

[0056] From the above description of the embodiments, those skilled in the art will clearly understand that the present invention can be implemented via hardware or via software combined with a necessary general-purpose hardware platform. Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product. This software product can be stored on a non-volatile storage medium (such as a CD-ROM, USB flash drive, or external hard drive) and includes instructions for enabling a computer device (such as a personal computer, server, or network device) to execute the methods described in the various embodiments of the present invention.

[0057] Those skilled in the art will appreciate that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes in the accompanying drawings are not necessarily required to implement the present invention.

[0058] Those skilled in the art will appreciate that the modules in the devices of the embodiments may be distributed in the devices of the embodiments as described in the embodiments, or may be located in one or more devices different from the embodiments with corresponding changes. The modules of the above embodiments may be combined into one module or further split into multiple submodules.

[0059] The serial numbers of the embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0060] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A method for generating a graphic report, characterized in that: include: Classifying the signal analysis data according to the file type and performing preprocessing to determine the image data and text data of the signal analysis data; Determining a first image processing model based on the image data, and processing the image data using the first image processing model to determine a first image feature vector; Determining a first text processing model based on the text data, and processing the first image feature vector and the text data using the first text processing model to generate a first text description; A retrieval-enhanced generation model is used to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data.

2. The method according to claim 1, wherein The determining of a first image processing model according to the image data comprises: classifying the image data according to image metadata of the image data or using a lightweight classification model to determine the image data type; A first image processing model is determined according to the image data type and a matching relationship between the image data type and the image processing model.

3. The method according to claim 1, wherein The method further includes: determining the first image processing model in the following manner: Read the performance headroom of the computing system; A first image processing model is determined according to the performance margin of the computing system.

4. The method according to claim 1, wherein The first image feature vector includes: target detection area attributes and anomaly detection area attributes; The target detection area attributes include: target detection area type, target detection area location, target detection area size and target detection area label; The anomaly detection area attributes include: anomaly detection area type, anomaly detection area location, anomaly detection area size and anomaly detection area label.

5. The method according to claim 1, wherein Determining a first text processing model according to the text data includes: Determining text processing model selection factors based on the text data; wherein the text processing model selection factors include: text type, text complexity and target output style; A first text processing model is determined according to the text processing model selection elements.

6. The method according to claim 1, wherein The retrieval-enhanced generation model is used to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data, including: Build a search knowledge base; Using the first image feature vector as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a first retrieval result; The first search result and the first text description are input into a first text processing model to generate a first graphic and text report of the signal analysis data.

7. The method according to claim 6, wherein The method further includes determining a top-k value in the first search result according to the image data type.

8. The method according to claim 6, wherein The method further comprises: Adjusting the correlation between the image and the text in the first graphic report to generate a second graphic report of the signal analysis data includes: Using an image understanding model to perform semantic back-matching on the first image-text report to determine a consistency assessment result; If the consistency assessment result is lower than a preset score, a semantic reasoning model is used to determine the missing text content in the first graphic report; Using the missing content of the text as retrieval information, the retrieval knowledge base is searched using a retrieval enhancement generation model to determine a second retrieval result; The second search result and the first graphic report are input into a first text processing model to generate a second graphic report of the signal analysis data.

9. A graphic report generating device, characterized in that: include: A preprocessing module, configured to classify the signal analysis data according to file type, and perform preprocessing to determine image data and text data of the signal analysis data; an image processing module, configured to determine a first image processing model based on the image data, and process the image data using the first image processing model to determine a first image feature vector; a text processing module, configured to determine a first text processing model based on the text data, and process the first image feature vector and the text data using the first text processing model to generate a first text description; A fusion module is used to adopt a retrieval enhancement generation model to fuse the first image feature vector and the first text description to generate a first graphic report of the signal analysis data.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the method for generating a graphic report according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Generation method and device of text processing model, equipment, medium and program product

    CN114861887A

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN115797717A

  • Image-text retrieval model construction method and system based on main semantic consistency

    CN117786143A

  • Medical quality control report intelligent generation method and system based on retrieval enhancement

    CN118571402A

  • Seal identification method and device, equipment, storage medium and product

    CN119068235A

Cited By

  • Image-text work generation method and device based on large model

    CN122220591A