Picture illustration method and device for long text content, electronic equipment and storage medium

Through the large natural language understanding model and multi-level linkage illustration strategy, the problem of insufficient fit between the illustrations of long articles and the full text is solved, efficient and accurate image generation and relevance evaluation are achieved, and the limitations of traditional keyword matching are broken through.

CN120723928APending Publication Date: 2025-09-30BEIJING ZERO ONE EVERYTHING INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510659860.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

Existing technologies have difficulty accurately grasping the overall intent when processing long articles, resulting in insufficient fit between the accompanying images and the full text. Traditional methods rely on keyword matching and are unable to understand deep semantics, contextual context, and emotional color.

Method used

A large natural language understanding model is used for semantic analysis, combined with a multi-level linkage image matching strategy, including local gallery retrieval, online gallery extended retrieval and image generation model, to generate adapted text descriptions or new images, and evaluate relevance through a multimodal model.

Benefits of technology

It improves the relevance and contextual fit of accompanying images, can accurately capture the core theme, emotional tendencies and implicit semantics of long texts, generate highly matching images, solve the semantic gap problem, and improve retrieval efficiency and success rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723928A_ABST
    Figure CN120723928A_ABST
Patent Text Reader

Abstract

The invention provides an illustration method and device for long text content, electronic equipment and a storage medium, and the illustration method for the long text content comprises the steps: inputting to-be-illustrated text content into a natural language understanding large model, and generating a content deep understanding result; executing a multi-level linkage illustration strategy based on a deep understanding result; performing correlation evaluation on the candidate images, and outputting an optimal illustration image according to an evaluation result; wherein the multi-level linkage illustration strategy comprises the following steps: performing a first round of retrieval in a local image library based on semantic vector similarity; when the first-round retrieval result does not meet the preset correlation threshold value, triggering expansion retrieval of the online image library; when the extended retrieval still does not meet the requirement, adaptive text description is generated, and an image generation model is called to create a new image, so that the core objective of a long text can be accurately captured, the semantic gap problem is solved, the retrieval range is gradually expanded or a high-matching-degree image is generated while the retrieval efficiency is ensured, and the retrieval efficiency is improved. And the success rate and the adaptability of illustration acquisition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a method, device, electronic device and storage medium for illustrating long text content. Background Art

[0002] With the explosive growth of internet information, digital content such as text, news, blogs, and social media posts is becoming increasingly abundant. To improve the readability, appeal, and efficiency of information dissemination, automatically matching high-quality, semantically relevant illustrations to this content has become a critical requirement. Accurate illustrations not only aid comprehension but also effectively capture user attention, enhancing the user experience. Traditional approaches to creating illustrations for multimodal digital content rely on keyword-based image matching within image libraries. However, these approaches rely heavily on literal keyword matching and struggle to understand the content's deeper semantics, context, abstract concepts, or emotional undertones. Related technologies include large-scale model-based image retrieval and text-to-image matching. These technologies can retrieve images from image libraries that are highly semantically relevant to a single sentence based on the semantic content of a short input text, improving the accuracy and relevance of illustrations for text content. While existing large-scale models have strong single-sentence understanding capabilities, they can still struggle to accurately grasp the overall intent when dealing with long articles or when illustrations require a comprehensive understanding of the overall text's theme, style, and logical relationships across multiple paragraphs. This can lead to quotations being taken out of context, resulting in an inadequate fit between the illustration and the full text. Summary of the Invention

[0003] The present invention provides a method, device, electronic device and storage medium for illustrating long text content, which are used to solve the defect that traditional content illustration methods have poor understanding ability of long text, resulting in insufficient fit between the illustration and the full text.

[0004] The present invention provides a method for adding pictures to long text content, comprising: Inputting the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; Execute a multi-level linkage image matching strategy based on the in-depth understanding result; Perform relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and output the optimal matching image based on the evaluation results; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

[0005] The method for providing a picture for a long text content provided by the present invention further includes: Obtaining image requirements, including an image target type; When the image target type is a picture, the generated content depth understanding result includes one or more keyword combinations for image retrieval and / or a descriptive text; When the target type of the illustration is a chart, the generated content depth understanding result includes data or logical relationships that can be charted.

[0006] According to the method for providing an image for long text content provided by the present invention, when the image matching target type is an image, executing a multi-level linkage image matching strategy based on the in-depth understanding result includes: Obtain image parameters, and filter out a subset of images that meet size requirements from the image library based on the image parameters; Performing a preliminary search in the filtered image subset using the one or more sets of keyword combinations for image retrieval and / or a descriptive text; When the initial search fails, performing an extended search in an online image search engine or a commercial image library, and filtering the extended search results using a multimodal model; When the extended search fails to output images that meet the requirements, the image generation model is called to deeply understand the text content of the image to be matched and generate a raw image description; The raw image description and the image parameters are input into an online image generation engine to generate an optimal image.

[0007] According to the method for providing images for long text content provided by the present invention, filtering the extended search results using a multimodal model includes: Calculating a cross-modal semantic relevance score between the expanded search result and the one or more keyword combinations for image retrieval and / or a descriptive text by the multimodal model; The expanded search results are filtered based on the cross-modal semantic relevance score.

[0008] According to the method for providing illustrations for long text content provided by the present invention, when the illustration target type is a chart, executing a multi-level linkage illustration strategy based on the in-depth understanding result includes: receiving a chart type selected by a user, and if the chart type is a fixed logic chart, filling the chartable data or logical relationship into a preset chart template; Generate an optimal chart based on the filled chart template.

[0009] According to the method for providing a diagram for long text content provided by the present invention, if the diagram type is an open diagram, the method includes: Calling a large model for drawing a chart, and generating a code snippet for drawing the chart based on the data or logical relationship available for charting; The code snippet is rendered in a browser environment, and the rendering result is captured as an optimal chart.

[0010] The method for providing a picture for a long text content provided by the present invention further includes: The new image and corresponding text metadata are stored in the local image library.

[0011] The present invention also provides a device for providing images for long text content, comprising: An input module is used to input the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; An execution module, configured to execute a multi-level linkage image matching strategy based on the in-depth understanding result; The output module is used to evaluate the relevance of the candidate images obtained after executing the multi-level linkage matching strategy, and output the optimal matching image based on the evaluation results; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

[0012] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method for illustrating long text content as described above is implemented.

[0013] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the above-described methods for illustrating long text content.

[0014] The present invention provides a method, device, electronic device and storage medium for matching long text content. The method inputs the text content to be matched into a natural language understanding large model, performs semantic analysis on the text content to be matched based on the natural language understanding large model, and generates a content deep understanding result; executes a multi-level linkage matching strategy based on the deep understanding result; performs relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and outputs the optimal matching according to the evaluation result; wherein the multi-level linkage matching strategy includes a first-round search based on semantic vector similarity in the local map library; when the first-round search result does not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling an image generation model to create a new image, breaking through the limitations of traditional keyword matching detection matching, and can accurately capture the core theme, emotional tendency and implicit semantics of long texts, solve the semantic gap problem, and adopt a multi-level linkage matching strategy. While ensuring retrieval efficiency, it gradually expands the search scope and even automatically generates high-matching images, thereby improving the success rate and adaptability of matching acquisition. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the present invention or the prior art, a brief introduction is given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is one of the flow charts of the method for providing pictures for long text content provided by the present invention; Figure 2 This is the second flow chart of the method for adding pictures to long text content provided by the present invention; Figure 3 This is the third flow chart of the method for providing pictures for long text content provided by the present invention; Figure 4 This is the fourth flow chart of the method for providing pictures for long text content provided by the present invention; Figure 5 This is the fifth flow chart of the method for providing illustrations for long text content provided by the present invention; Figure 6 This is the sixth flow chart of the method for providing illustrations for long text content provided by the present invention; Figure 7 It is a structural schematic diagram of the device for providing illustrations for long text content provided by the present invention. DETAILED DESCRIPTION

[0017] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0018] Figure 1 A flowchart of a method for providing a diagram for long text content provided by an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method for adding pictures to long text content provided by the embodiment of the present invention includes: Step 101: Input the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; Step 102: executing a multi-level linkage image matching strategy based on the depth understanding result; Step 103: perform a correlation evaluation on the candidate images obtained after executing the multi-level linkage image matching strategy, and output the optimal image matching according to the evaluation result; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

[0019] Although traditional large-model-based image retrieval and image-text matching technology can retrieve images that are highly semantically relevant to a single sentence in the image library based on the semantic content of the input short text, it may still be difficult to accurately grasp the overall intention when processing long articles or when it is necessary to integrate the full text's theme, style, and multi-paragraph logical relationships to create illustrations. It is easy to "take things out of context", resulting in insufficient fit between the illustration and the full text.

[0020] The embodiment of the present invention provides a method for matching long text content with pictures. The method inputs the text content to be matched with pictures into a natural language understanding large model, performs semantic analysis on the text content to be matched with pictures based on the natural language understanding large model, and generates a deep understanding result of the content; executes a multi-level linkage matching strategy based on the deep understanding result; performs relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and outputs the optimal matching picture according to the evaluation result; wherein the multi-level linkage matching strategy includes a first-round search based on semantic vector similarity in the local map library; when the first-round search result does not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling an image generation model to create a new image, breaking through the limitations of traditional keyword matching detection matching, and can accurately capture the core theme, emotional tendency and implicit semantics of long texts, solve the semantic gap problem, and adopt a multi-level linkage matching strategy. While ensuring retrieval efficiency, it gradually expands the search scope and even automatically generates high-matching images, thereby improving the success rate and adaptability of matching pictures.

[0021] Based on any of the above embodiments, Figure 2 As shown, in an embodiment of the present invention, it also includes: Step 201: Obtain image matching requirements, including image matching target types; Step 202: When the image target type is a picture, the generated content depth understanding result includes one or more keyword combinations for image retrieval and / or a descriptive text; Step 203: When the target type of the image is a chart, the generated content depth understanding result includes data or logical relationships that can be charted.

[0022] In an embodiment of the present invention, the system receives user input of the content to be illustrated (e.g., a text paragraph or article), optional contextual information (e.g., the overall theme of the article, the context), and clear illustration requirements. Illustration requirements include: target type, i.e., whether an image or a chart is required. If the target type is an image, illustration requirements may include, for example, the desired image size (or aspect ratio), primary color, style (e.g., photography, cartoon, painting, etc.), and elements to avoid (e.g., specific colors, sensitive content classifications). If the target type is a chart, illustration requirements may include, for example, the desired chart type (e.g., bar chart, line chart, pie chart), data point prompts, etc.

[0023] Leveraging natural language understanding capabilities, the language model conducts in-depth semantic analysis of the input content and context, understanding its core themes, key information, emotional tone, implicit meaning, and style. Simultaneously, it analyzes the user's input image requirements and uses them as important constraints for subsequent retrieval and generation.

[0024] Based on an understanding of the content and requirements, we leverage natural language understanding capabilities to determine the overall direction and characteristics of the required illustrations. For image requirements, we generate one or more keyword combinations (search tags) and / or a descriptive text (search description) for image retrieval. This description strives to capture the essence of the content and incorporate user needs (e.g., style and atmosphere). For chart requirements, we initially determine whether the content contains data or logical relationships suitable for charting.

[0025] Based on any of the above embodiments, Figure 3 As shown, when the image matching target type is a picture, executing the multi-level linkage image matching strategy based on the depth understanding result includes: Step 301: Obtain image parameters, and select a subset of images that meet size requirements from the image library based on the image parameters; The embodiment of the present invention fully meets the user's needs for practicality and customization of illustrations by obtaining image parameters, thereby improving ease of use. Receive user requirements for specific attributes such as image / chart size, color, style, type, and always process and filter these practical requirements as core constraints in subsequent internal retrieval, online retrieval, online image generation, and chart generation. Combined with the content security filtering mechanism, it ensures that the output illustrations are not only semantically relevant, but also clear, watermark-free (or processed on demand), size-compliant, and content-safe. This greatly reduces the workload of users who still need to perform a large amount of manual screening, editing, and adjustment after obtaining preliminary results, significantly improves the direct usability and user satisfaction of illustrations, and solves the problems of uneven quality and poor convenience of keyword search results in the background technology.

[0026] Step 302: Perform a preliminary search in the filtered image subset using the one or more sets of keyword combinations for image retrieval and / or a descriptive text; Step 303: When the initial search fails, perform an extended search in an online image search engine or a commercial image library, and filter the extended search results using a multimodal model; Step 304: When the extended search fails to output a picture that meets the requirements, a large picture generation model is called to deeply understand the content of the text to be matched with the picture and generate a raw picture description; Step 305: Input the raw image description and the image parameters into an online image generation engine to generate an optimal image.

[0027] This embodiment of the present invention utilizes a multi-stage, multi-level linkage strategy (internal image library -> online search -> online image generation) to significantly expand the sources of available images. In particular, the optimized generation mechanism for image generation instructions enables the creative generation of new, high-quality images that meet specific requirements when exhaustive search methods still fail to find suitable images. This breaks through the limitations of traditional image libraries and provides unprecedented visual expression possibilities for content creation.

[0028] The multi-stage, multi-level linkage strategy adopted in the embodiment of the present invention specifically includes: Internal gallery retrieval stage: Based on the image size required by the user or the recommended size range preset by the system based on the content layout, first screen out a subset of images that meet the size requirements in the internal gallery. In the screened image subset, use the search tags generated in step one to perform a preliminary search through traditional keyword matching or vector similarity retrieval methods. For the images initially retrieved, use the first scoring mechanism, such as a lightweight semantic similarity model, to calculate their relevance scores with the search tags / descriptions. Set a relevance score threshold T1. If a picture scores higher than T1, it is considered to be coarse-grained related. The system determines whether it is necessary to further refine the sorting of coarse-grained related pictures based on preset strategies or user requirements. For example, if there are a large number of coarse-grained related pictures, or if the user has extremely high requirements for the semantic fit between the pictures and the content, post-processing is initiated.

[0029] In some embodiments of the present invention, the filtering of the extended retrieval results through the multimodal model includes: calculating the cross-modal semantic relevance score between the extended retrieval results and the one or more groups of keyword combinations for image retrieval and / or a descriptive text through the multimodal model; and filtering the extended retrieval results based on the cross-modal semantic relevance score.

[0030] This embodiment of the present invention uses a multimodal model to perform secondary sorting and screening of search results. This multimodal model calculates cross-modal semantic similarity between the text of the input content and the visual content of the candidate images, generating a more accurate relevance score. Based on this score, secondary sorting is performed to select the optimal image or images. If an image that meets all requirements is found at this stage, it is output as the final result. If no images that meet the requirements are found, or if the optimal image's relevance score remains below a preset higher threshold T2, the online search phase automatically begins.

[0031] Online image retrieval stage: Figure 4As shown, in some cases, such as when the content is novel, requires unique details, or covers niche areas, the local image library cannot match the optimal image. In this case, the system automatically adapts and populates the API search parameters of mainstream online image search engines or commercial image libraries based on the search tags and user-specified image attributes (such as size, color, and format) to execute the online search request. The initial results returned by the online search undergo a first round of rapid filtering based on the user's specified size and image quality, for example, excluding images with low resolution or obvious blur. For images that pass this initial filtering, the multimodal model is invoked again. A cross-modal semantic relevance score is calculated for each image with the input content, and the images are ranked accordingly. During this process, the multimodal model or an auxiliary content security moderation model simultaneously analyzes the image content and automatically excludes images containing pre-set filtering factors (e.g., violence, pornography, horror, or inappropriate content as defined by laws, regulations, or platform policies).

[0032] Select one or more images that are optimal after sorting and pass the content filter. The system will perform structured processing on the selected online images and store them in the internal gallery. Save the image file and record the image metadata (such as source, original URL, size, format, etc.). Use the natural language understanding model or image description model to generate new, richer labels and semantic description text for the image. Record the original input content or its summary of this pairing, as well as the relevance score calculated by the multimodal model. Enrich the internal gallery and provide higher quality and faster matching for subsequent searches of the same or similar content, thereby reducing dependence on online searches and improving the overall efficiency and effectiveness of the system. If a picture that meets the requirements is found at this stage, it will be output as the final result. If no suitable picture is found (for example, the relevance of the online search results is not high, or none of them pass the content filter), it will automatically enter the online image generation stage.

[0033] Online image generation stage: Figure 5 As shown in the figure, when the search fails to meet the requirements, the system determines that image generation is necessary. At this point, the large image generation model is called again to gain a deeper understanding and creative refinement of the original input content. The goal of the large image generation model is to generate a clear, unambiguous, detailed, generalizable, and core semantic description (prompt). This description will specifically consider the user-specified size requirements and may incorporate experience from previous search failures (for example, avoiding certain elements that lead to low relevance). If the content involves a specific field (such as medical or technology), the model will attempt to use specialized terminology or style cues in that field.

[0034] Submit the optimized raw image description and user-specified size parameters to one or more online image generation engines / models (e.g., services based on technologies such as Stable Diffusion and DALL-E). Receive the generated image. The system may perform multiple attempts or fine-tune the generation parameters as needed to achieve the optimal image.

[0035] In this embodiment of the present invention, the multimodal model can be used again to assess the relevance of the generated image to the original content, ensuring that the generated image meets expectations. Content security audits are also conducted. The final selected generated image is similarly structured and stored in an internal image library, accumulating high-quality generated assets and providing reference or direct reuse for similar future image generation needs. The generated image that meets the requirements is output as the final result.

[0036] The embodiments of the present invention significantly improve the relevance and contextual fit of the accompanying images. By deeply understanding the input content, optional context, and the user's explicit needs, and using a multimodal model to perform cross-modal semantic alignment of text and image / chart visual content, the core theme, subtle emotions, logical relationships, and overall style of the content can be more accurately grasped. This effectively overcomes the semantic gap problem that exists in keyword-based matching methods, as well as the defects of some existing large-scale model applications that may appear out of context and lack understanding of the overall intent when processing long texts or complex contexts. It can provide accompanying images that are highly consistent with the original text in terms of semantics, style, and emotion, thereby enhancing the expressiveness and appeal of the content.

[0037] In the embodiment of the present invention, Figure 6 As shown, when the target type of the image is a chart, executing the multi-level linkage image strategy based on the deep understanding result includes: Step 401: receiving a chart type selected by a user. If the chart type is a fixed logic chart, filling the chartable data or logical relationship into a preset chart template. Step 402: Generate an optimal chart based on the filled chart template.

[0038] In an embodiment of the present invention, if the chart type is an open chart, it includes: Step 501: Calling a large model for drawing a chart, and generating a code snippet for drawing the chart based on the data or logical relationship available for charting; Step 502: Render the code snippet in a browser environment, and capture the rendering result as an optimal chart.

[0039] In an embodiment of the present invention, a large charting model analyzes input content, attempting to identify and extract structured or semi-structured data, key indicators, trends, comparative relationships, and other information suitable for presentation in a chart. Based on the user-specified chart type (if any) or the model's recommended chart type based on extracted data features, the system enters the corresponding generation logic. If a common chart type (such as a bar chart, line chart, or pie chart) is selected, the system populates the extracted structured data into a pre-set chart template. Chart templates have adaptive capabilities, such as automatically adjusting axis ranges and label display based on data volume. The resulting chart template format is consistent with the interactive display page. If a more flexible or customized chart is required, or if the content data does not fully conform to a fixed template, the system can leverage the capabilities of the large charting model to directly generate the HTML, CSS, and JavaScript code snippets required to draw the chart based on the extracted data and the desired chart presentation logic. The generated code snippets are then rendered in a headless browser environment, and the rendered results are captured as images. The generated chart images are then output as the final result.

[0040] In an embodiment of the present invention, the existing technology for charting requirements suffers from the narrow applicability and poor flexibility of template or rule-based methods. This embodiment of the present invention utilizes a large charting model to perform in-depth content analysis to extract structured data. This dual-path mechanism, combining fixed-logic chart template filling with innovative open chart generation (e.g., HTML / code), enables the automatic generation of accurate and flexible charts for diverse text content. In particular, the open chart generation capability allows for the generation of highly customized chart code based on data characteristics and presentation requirements, greatly enhancing the power and convenience of content data visualization and overcoming the limitations of traditional methods.

[0041] In an embodiment of the present invention, the method for providing images for long text content further includes: storing the new image and corresponding text metadata into the local image library.

[0042] In the embodiment of the present invention, the text metadata includes tags, image descriptions, image size, classification and other data.

[0043] In this embodiment of the present invention, images generated through large-scale model processing generate rich labels and semantic descriptions, which are then stored in an internal image library. This not only dynamically expands and optimizes the quality and coverage of the internal image library, but also enables the system to learn from past successful cases. For subsequent image requests for the same or similar content, the system can prioritize rapid matching from the higher-quality internal image library, thereby significantly improving retrieval efficiency and reducing reliance on online retrieval and image generation, which consume large amounts of computing resources and have long response times. In the long run, this helps reduce operating costs and improve system response speed.

[0044] The reasoning calculation of existing large models (especially dense retrieval or when a large number of candidate images need to be scored) usually requires large computing resources (such as GPUs). In online application scenarios that require high concurrency and low-latency responses, efficiency and cost may become bottlenecks. In the process of processing requests for retrieving images, the embodiment of the present invention will give priority to disassembling the requests, constraining the search scope in advance, and reducing the search cost. The images retrieved by the retrieval will be filtered twice to reduce the number entering the multimodal model and reduce computing consumption. The calculated and generated images are structurally stored and enter the local gallery retrieval system. For subsequent image requests for the same or similar content, the system can give priority to quick matching from the higher-quality internal gallery, improve retrieval efficiency, reduce repeated calculation processes, and quickly return results. In the long run, it will help reduce operating costs.

[0045] Existing image retrieval technology based on large models still has shortcomings. For example, while existing models have strong single-sentence comprehension capabilities, they may still struggle to accurately grasp the overall intent when dealing with long articles or when integrating the main theme, style, and logical relationships of multiple paragraphs. This can lead to quotations being taken out of context, resulting in an image that is not fully aligned with the full text. Images must not only be relevant to the content but also aesthetically pleasing in terms of size, clarity, and style. Retrieval technology that relies solely on large models cannot return satisfactory results. Large models can, to a certain extent, address the semantic gap encountered by traditional keyword matching. However, this relies on the availability of valid image sources, which cannot be addressed by relying solely on large models. For applications such as medical, legal, and specialized industry reports, a lack of domain knowledge can lead to misunderstandings and the retrieval of unprofessional images.

[0046] This embodiment of the present invention integrates three phases: internal image library search, online image search, and online image generation. It makes intelligent decisions and streamlines the flow of results based on user needs and the results of each phase, achieving a balance between efficiency, cost, and image quality. In the initial image matching phase, a large model is used to jointly analyze content, context, and user-specified image / chart attributes (such as size, color, and type) to generate more accurate initial search information or image generation instructions. If the search yields no results, the original text is not simply reused. Instead, the large model is used to reinterpret and creatively refine the content, generating a detailed, optimized, and compliant raw image description that is more suitable for the image generation model. High-quality images obtained through online search and online generated images are processed by the large model (generating labels and descriptions) and stored in the internal image library, enabling dynamic growth and self-optimization of the library, improving the efficiency and relevance of subsequent searches. The online search and optional raw image evaluation phases closely integrate multimodal semantic relevance ranking with content safety filtering (such as the identification of pornographic, violent, and terrorist content) to ensure the relevance and compliance of the output results. To address charting needs, we not only support fixed templates based on extracted data but also innovatively introduce open chart generation methods that leverage large models to directly generate HTML and other drawing logic based on content and data characteristics, and then render the resulting charts. User-specified illustration requirements are considered and incorporated throughout multiple steps of the entire process (such as internal library size screening, online search parameter adaptation, raw image size specification, and chart type selection), enhancing the practicality and satisfaction of the final results.

[0047] The method for illustrating long text content provided by the embodiments of the present invention can deeply understand the semantics and context of the content, comprehensively consider practical requirements (size, quality, style, etc.), and has the capabilities of multi-source retrieval (internal libraries, online libraries) and on-demand generation (images, charts), and can be processed automatically, thereby significantly improving the quality, efficiency, and user satisfaction of content illustrations. Users only need to enter the content and basic requirements to obtain high-quality illustrations that meet multi-dimensional needs, which significantly reduces manual intervention, shortens the content production and release cycle, and improves overall work efficiency. It effectively solves many pain points of existing content illustration technology in terms of relevance, practicality, coverage, flexibility, and efficiency, and provides a more powerful and intelligent automated illustration solution for various digital content platforms.

[0048] The following describes the device for matching pictures of long text content provided by the present invention. The device for matching pictures of long text content described below and the method for matching pictures of long text content described above can refer to each other.

[0049] Figure 7 A schematic diagram of the structure of a device for providing a diagram for a long text content according to an embodiment of the present invention is shown in FIG. Figure 7 As shown, the device for providing a picture of a long text content provided by an embodiment of the present invention includes: Input module 701, for inputting the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; An execution module 702 is configured to execute a multi-level linkage image matching strategy based on the depth understanding result; Output module 703, used to perform relevance evaluation on the candidate images obtained after executing the multi-level linkage image matching strategy, and output the optimal image matching according to the evaluation result; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

[0050] The device for matching long text content provided by the embodiment of the present invention inputs the text content to be matched with an image into a natural language understanding large model, performs semantic analysis on the text content to be matched with an image based on the natural language understanding large model, and generates a content deep understanding result; executes a multi-level linkage matching strategy based on the deep understanding result; performs relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and outputs the optimal matching image according to the evaluation result; wherein the multi-level linkage matching strategy includes a first-round search based on semantic vector similarity in the local map library; when the first-round search result does not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling an image generation model to create a new image, breaking through the limitations of traditional keyword matching detection matching, and can accurately capture the core theme, emotional tendency and implicit semantics of long texts, solve the semantic gap problem, and adopt a multi-level linkage matching strategy. While ensuring retrieval efficiency, it gradually expands the search scope and even automatically generates high-matching images, thereby improving the success rate and adaptability of matching image acquisition.

[0051] Based on any of the above embodiments, this embodiment provides a method for automatically matching images to presentation content (such as Microsoft PowerPoint, Google Slides, etc.), and is applied to an intelligent image matching system. In this embodiment, the system intelligently matches or generates appropriate images or charts for a specific presentation page based on the presentation's content, contextual information, and style attributes (such as theme and layout). The intelligent image matching system can be a standalone application or a plug-in / module integrated into presentation editing software. The system primarily includes the following functional modules: Input receiving module: Responsible for receiving presentation data input by the user, including the text content of a single page, the overall context information of the presentation (such as the title, chapter structure, and the content of the previous and next pages), and the illustration requirements parsed from the presentation template / layout (such as the desired image / chart size and style).

[0052] Content Analysis and Strategy Determination Module: Integrates the first major model to deeply understand the input text content and context, parse user / system image requirements, and preliminarily determine the required image type (picture or chart) and preliminary retrieval / generation direction.

[0053] Image Processing Module: This module integrates a multi-stage image retrieval and generation mechanism, including an internal image library retrieval submodule, an online image retrieval submodule, and an online image generation submodule. This module collaboratively utilizes the second-largest model (multimodal model), the third-largest model (image generation instruction optimization model), and the security content review model.

[0054] Chart processing module: Integrates chart generation logic and synergistically utilizes the fourth model (data extraction and code generation model).

[0055] Internal image library and knowledge base: Store local image resources and processed (such as annotation and description) online acquired and generated images as priority retrieval sources and achieve knowledge enhancement.

[0056] Output management module: responsible for outputting the finalized illustration results (pictures or charts) to the designated location of the presentation.

[0057] User interface module: provides a user interaction interface, allowing users to input and adjust requirements, or provide feedback on the illustration results.

[0058] The system can run on a user's local computing device (e.g., a personal computer equipped with a processor, memory, storage, and network interface), or it can be deployed as a cloud service on a remote server, interacting with the user's device over the network. Internal image libraries and various large models can be stored in local memory or on a cloud server, and various functions are completed by the processor executing instructions stored in memory.

[0059] In this embodiment, the process of the method for intelligently matching PPT content with images based on a large model is as follows: (1) The input receiving module receives the text content of the current PPT page, the title of the entire PPT, chapter information, previous and next page text, and other context, as well as information parsed from the PPT template, theme, and layout, such as the image placeholder size preset in the page layout (which determines the expected image size or aspect ratio), the main color tone and overall visual style defined by the PPT theme (such as business, technology, art, etc.). Users can also enter more specific image requirements through the user interface module, such as specifying specific keywords or excluding certain elements.

[0060] (2) The content analysis and strategy determination module uses the first model to perform in-depth semantic analysis on the received PPT page text and context. For example, it analyzes whether the page discusses team collaboration, market trends, or technical details. At the same time, it analyzes the image requirements extracted from the PPT attributes and user input (such as size 640x480, "business flat" style, avoid red tones, etc.).

[0061] (3) Based on the understanding of the PPT content and requirements, the first model preliminarily determines whether a picture is needed to visualize the concept or a chart is needed to display the data. If it is determined to be a picture requirement, the model will generate a set of keyword combinations for the PPT page content (for example, "teamwork", "collaborative office", "meeting") and / or a descriptive text (for example, "a business-style picture showing team members actively discussing in a meeting room, with the main color being blue"). This information will serve as an important basis for subsequent image retrieval / generation and incorporate PPT / user requirements such as size, style, and color. If it is determined to be a chart requirement, the model will further analyze whether there is structured data or logical relationships in the PPT page text.

[0062] (4) The internal image library retrieval submodule first selects a subset of images with a size range that meets the requirements from the internal image library based on the size requirements obtained from the PPT layout. In the selected subset, relevant images are preliminarily searched through keyword matching or vector similarity retrieval. The first scoring mechanism is used to calculate the relevance score of the retrieved images with the retrieval information output in step 1. A threshold T1 is set (for example, 0.7). If there are images with a score higher than T1, they are considered preliminarily relevant. The system determines whether further refinement is required based on a preset strategy (for example, if there are a large number of high-scoring images, or if the PPT content has extremely high requirements for image semantic matching). If it is decided to perform refinement, the second largest model (multimodal model) is called to perform cross-modal semantic similarity calculations on the PPT page text and the preliminarily retrieved images to obtain a more accurate relevance score. A higher threshold T2 is set (for example, 0.9). It is determined whether the score of the optimal image is higher than T2 and meets all requirements (for example, it passes the preset style matching check and does not contain user-specified avoidance elements). If the conditions are met, the image is output to the designated location on the PPT page through the output management module (106), and the process ends. If the conditions are not met, or no preliminary relevant images are found in the internal image library, the process automatically enters the online image retrieval stage.

[0063] (5) The online image retrieval submodule adapts the API parameters of mainstream online image search engines or commercial image libraries based on the search tags and user-specified image attributes (such as size, style, and tone) to execute online search requests. The preliminary results returned by the online search are first quickly filtered based on the size requirements obtained from the PPT layout and basic image quality standards to exclude images that do not meet the basic conditions.

[0064] (6) For images that have passed the initial filtering, the second largest model (multimodal model) is called again to calculate the cross-modal semantic relevance score between the image and the PPT page text for sorting. Key step: At this stage, the security content review model is collaboratively called to simultaneously perform content analysis on each image, automatically excluding any images containing inappropriate content such as violence, pornography, and terror (as defined by the platform or laws and regulations), ensuring that the output images are safe and compliant.

[0065] (7) Select one or more images that are optimal after sorting and pass the content filter. The selected online images are structurally processed and stored in the internal image library. Specifically, it includes: saving the image file; recording metadata such as the source URL and size; using the first model or a special image description model to generate new, richer tags and semantic description text for the image for subsequent retrieval; recording the content summary and relevance score of the PPT page of this pairing, and building the association between content and images. This step continuously enriches the internal image library and improves the efficiency and quality of future image matching. Determine whether an image that meets the requirements and passes the security filter is found. If so, the image is output to the PPT through the output management module, and the process ends. If no suitable image is found (for example, the relevance of the online search results is generally not high, or none of them pass the content filter), it automatically enters the online image generation stage.

[0066] (8) When the retrieval fails to meet the requirements, the online image generation submodule determines that image generation is required. Key steps: At this time, the third model (raw image instruction optimization model) is called. This model conducts a deeper understanding and creative refinement of the original PPT page text, while taking into account the size, style and other requirements obtained from the PPT attributes and the experience of failure in the first two retrieval stages (for example, avoiding certain elements that lead to low relevance), and generates a clear, detailed and PPT image description (Prompt) that meets the PPT illustration requirements. The optimized raw image description and size parameters are submitted to one or more online raw image engines / models (for example, using API services based on technologies such as Stable Diffusion and DALL-E) for image generation. The system can make multiple attempts or fine-tune the generation parameters as needed to achieve better results. The second model (multimodal model) is used again to evaluate the relevance of the generated image to the PPT page content and conduct a content security audit to ensure that the generated image meets expectations and is safe. The final selected generated image is also structured and stored in the internal image library. The steps are the same as the knowledge enhancement processing in the online retrieval stage. This helps accumulate high-quality generated assets for reference or direct reuse for future illustration needs of similar PPT content.

[0067] (9) The generated images that meet the requirements are output to the designated location on the PPT through the output management module, and the process ends. The chart processing module is started according to the strategy judgment result: the fourth model (data extraction and code generation model) is called to analyze the PPT page text, and try to identify and extract the numbers, indicators, trend descriptions, comparative relationships, and other information suitable for display in charts. According to the chart type specified by the user (for example, the user requires "display with a bar chart") or the chart type recommended by the model based on the extracted data features, the system enters the corresponding generation logic: Fixed-logic chart generation: If you select a common chart type (such as a bar chart, line chart, or pie chart), the system will extract the structured data and populate it into a pre-set chart template. The template will then adapt to the data volume and PPT layout (size), producing a chart template format that is suitable for PPT presentation and interaction.

[0068] Open Chart Generation (HTML / Code Drawing): Key Steps: If a more flexible, customized, or complex chart is required, or if the extracted data doesn't fully conform to a fixed template, the system leverages the power of the fourth model to directly generate the HTML, CSS, and JavaScript code snippets used to draw the chart based on the extracted data and the desired chart presentation logic (for example, using the call logic of a charting library like D3.js or ECharts.js). The system then executes these code snippets in a headless browser environment, renders the chart, and captures the rendered result as an image.

[0069] (10) The generated chart image is output to the specified location on the PPT through the output management module, and the process ends.

[0070] If after all the above stages the system still cannot find or generate an image that meets all requirements (including security filtering), the system may prompt the user through the user interface module or provide alternative solutions.

[0071] The system described in the present invention can be implemented using standard computing devices, such as a server cluster or a user's personal computer. The functions of the processing module, internal image library, and output management module are performed by one or more processors executing computer program instructions stored in memory. The processor communicates with external resources (such as online image libraries, online image generation engines, and cloud-based large model services) through a network interface. The internal image library data can be stored on a local hard disk or in a remote database. Various large models can be deployed on a local computing device or call a remote cloud-based model service through an API. The user interface module interacts with the user through a display and input devices. In an embodiment of the present invention, the internal image library and related models can be stored in a local memory or in a cloud server.

[0072] An embodiment of the present invention further provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus. The processor may call logic instructions in the memory to execute a method for matching long text content with images, the method comprising: inputting the text content to be matched with images into a large natural language understanding model, performing semantic analysis on the text content to be matched with images based on the large natural language understanding model, and generating a content deep understanding result; executing a multi-level linkage matching strategy based on the deep understanding result; performing a relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and outputting an optimal matching image based on the evaluation result; wherein the multi-level linkage matching strategy includes performing a first-round search based on semantic vector similarity in the local image library; when the first-round search results do not meet a preset relevance threshold, triggering an extended search of the online image library; when the extended search still does not meet the requirements, generating an adapted text description and calling an image generation model to create a new image.

[0073] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for illustrating long text content provided by the above-mentioned methods, the method comprising: inputting the text content to be illustrated into a large natural language understanding model, performing semantic analysis on the text content to be illustrated based on the large natural language understanding model, and generating a content deep understanding result; executing a multi-level linkage illustration strategy based on the deep understanding result; performing a relevance evaluation on the candidate images obtained after executing the multi-level linkage illustration strategy, and outputting the optimal illustration based on the evaluation result; wherein, the multi-level linkage illustration strategy includes a first-round search based on semantic vector similarity in the local map library; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

[0074] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0075] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A method for arranging pictures for long text content, characterized in that: include: Inputting the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; Execute a multi-level linkage image matching strategy based on the in-depth understanding result; Perform relevance evaluation on the candidate images obtained after executing the multi-level linkage matching strategy, and output the optimal matching image based on the evaluation results; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

2. The method for arranging images for long text content according to claim 1, characterized in that: Also includes: Obtaining image requirements, including an image target type; When the image target type is a picture, the generated content depth understanding result includes one or more keyword combinations for image retrieval and / or a descriptive text; When the target type of the illustration is a chart, the generated content depth understanding result includes data or logical relationships that can be charted.

3. The method for arranging pictures for long text content according to claim 2, characterized in that: When the image matching target type is a picture, executing the multi-level linkage image matching strategy based on the depth understanding result includes: Obtain image parameters, and filter out a subset of images that meet size requirements from the image library based on the image parameters; Performing a preliminary search in the filtered image subset using the one or more sets of keyword combinations for image retrieval and / or a descriptive text; When the initial search fails, performing an extended search in an online image search engine or a commercial image library, and filtering the extended search results using a multimodal model; When the extended search fails to output images that meet the requirements, the image generation model is called to deeply understand the text content of the image to be matched and generate a raw image description; The raw image description and the image parameters are input into an online image generation engine to generate an optimal image.

4. The method for arranging pictures for long text content according to claim 3, characterized in that: The filtering of the extended search results by the multimodal model includes: Calculating a cross-modal semantic relevance score between the expanded search result and the one or more keyword combinations for image retrieval and / or a descriptive text by the multimodal model; The expanded search results are filtered based on the cross-modal semantic relevance score.

5. The method for arranging pictures for long text content according to claim 2, characterized in that: When the target type of the image is a chart, executing the multi-level linkage image strategy based on the in-depth understanding result includes: receiving a chart type selected by a user, and if the chart type is a fixed logic chart, filling the chartable data or logical relationship into a preset chart template; Generate an optimal chart based on the filled chart template.

6. The method for arranging pictures for long text content according to claim 5, characterized in that: If the chart type is an open chart, it includes: Calling a large model for drawing a chart, and generating a code snippet for drawing the chart based on the data or logical relationship available for charting; The code snippet is rendered in a browser environment, and the rendering result is captured as an optimal chart.

7. The method for arranging pictures for long text content according to claim 1, characterized in that: Also includes: The new image and corresponding text metadata are stored in the local image library.

8. A device for providing pictures for long text content, characterized in that: include: An input module is used to input the text content to be matched with an image into a natural language understanding model, so as to perform semantic analysis on the text content to be matched with an image based on the natural language understanding model and generate a content in-depth understanding result; An execution module, configured to execute a multi-level linkage image matching strategy based on the in-depth understanding result; The output module is used to evaluate the relevance of the candidate images obtained after executing the multi-level linkage matching strategy, and output the optimal matching image based on the evaluation results; Among them, the multi-level linkage illustration strategy includes performing a first-round search in the local map library based on semantic vector similarity; when the first-round search results do not meet the preset relevance threshold, triggering an extended search of the online map library; when the extended search still does not meet the requirements, generating an adapted text description and calling the image generation model to create a new image.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method for illustrating long text content according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for illustrating long text content according to any one of claims 1 to 7 is implemented.