An AI-based article illustration method, computer-readable medium, and device

Through large-model deep learning technology and intelligent image matching methods, combined with text-generated image and image-generated image models, the low efficiency, high cost and copyright issues of traditional article image matching have been solved, and efficient and accurate article image matching has been achieved, improving content quality and visual experience.

CN120451340BActive Publication Date: 2025-09-19苏州日报社
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510960889.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-12
Publication Date
2025-09-19
Estimated Expiration
2045-07-12

AI Technical Summary

Technical Problem

The traditional article illustration process is inefficient, costly, and inaccurate. In particular, the issue of image copyright ownership plagues content creators and media organizations.

Method used

Large-scale deep learning technology is used to accurately extract text summaries and keywords, and large-scale models of text-to-image and image-to-image are combined for intelligent image generation. The CLIP model is used to review the semantic matching between images and texts, and multiple rounds of image generation are achieved.

Benefits of technology

It improves the fit between the illustrations and the article theme, enhances the content quality, solves the problems of low illustration efficiency and copyright conflicts, and enriches the readers' visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451340B_ABST
    Figure CN120451340B_ABST
Patent Text Reader

Abstract

The present invention discloses an article AI illustration method based on artificial intelligence, a computer-readable medium, and a device. The method uses preset prompt words for extracting text summaries and keywords, and extracts text summaries and keyword lists from selected text content through a large language model. Determine whether to use a private gallery, and retrieve a list of illustration pictures from the private gallery based on the extracted text summary and keyword list. Determine whether to use pictures directly; if so, insert the illustration pictures directly into the corresponding position of the text content to be illustrated. If not, continue to determine whether there are reference pictures, and regenerate the illustration picture list based on the raw picture large model, and determine whether there are suitable pictures in the regenerated illustration picture list. If so, insert the illustration pictures directly into the corresponding article of the text content to be illustrated, and end the article intelligent illustration process. If not, generate an illustration picture list until there is a suitable picture and insert it into the corresponding position of the text content to be illustrated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to an artificial intelligence-based article AI illustration method, computer-readable medium, and device. Background Art

[0002] Driven by the rapid development of mobile internet and large language models, the media industry is undergoing an unprecedented transformation. In this era of overwhelming information, how to stand out from the vast amount of content, capture readers' attention, and enhance the overall reading experience has become a core issue that all major media platforms are vying to address. Against this backdrop, article illustrations are becoming increasingly important as a key strategy for enhancing content appeal and increasing user engagement.

[0003] However, the traditional article illustration process faces many challenges, including but not limited to: low illustration efficiency, high cost, unsatisfactory accuracy, etc., especially the issue of image copyright ownership, which is a headache for many content creators and media organizations. These problems have always plagued content creators and media organizations. Summary of the Invention

[0004] The following is a brief summary of one or more aspects to provide a basic understanding of these aspects. This summary is not an exhaustive overview of all conceivable aspects and is neither intended to identify key or critical elements of all aspects nor to define the scope of any or all aspects. Its sole purpose is to present some concepts of one or more aspects in a simplified form as a prelude to the more detailed description that will be provided later.

[0005] The purpose of the present invention is to solve the above problems and provide an AI-based article illustration method, computer-readable medium, and device, which use large-model deep learning technology to accurately extract text summaries and keywords, and then perform intelligent illustration based on the extracted text summaries and keywords, thereby subverting the traditional illustration model.

[0006] The technical solution of the present invention is:

[0007] The present invention provides an AI-based article picture matching method, comprising the following steps:

[0008] Step S1: Select the text content to be matched with the image, and input the selected text content into the preset language extraction model to extract the text summary and keyword list;

[0009] Step S2: Determine whether to use a private image library; if so, retrieve a list of accompanying images from the private image library based on the extracted text summary and keyword list; if not, generate a list of accompanying images based on a preset raw image model;

[0010] Step S3: Based on the generated picture list, determine whether to directly use the picture; if so, directly insert the picture at the corresponding position of the text content to be illustrated to complete the article intelligent picture matching process; if not, determine whether there is a reference picture, and if so, regenerate the picture list based on the reference picture using the raw picture model; if not, return to step S2 to generate the picture list based on the preset raw picture model;

[0011] Step S4: Based on the regenerated picture list, determine whether there is a suitable picture; if so, directly insert the picture into the corresponding article of the text content to be illustrated, and end the article intelligent picture matching process; if not, continue to generate the picture list based on the regenerated picture list until a suitable picture is found and inserted into the corresponding position of the text content to be illustrated, and complete the article intelligent picture matching process

[0012] Among them, in step S2, the raw image large model includes a text-generated image large model and a picture-generated image large model; wherein, the artificial intelligence-based article AI illustration method determines that the private image library is not currently in use, generates illustration generation prompt words based on the extracted text summary and keyword list, and then uses the text-generated image large model or the picture-generated image large model to perform intelligent illustration generation based on the illustration generation prompt words, thereby obtaining an illustration picture list.

[0013] According to one embodiment of the artificial intelligence-based article AI illustration method of the present invention, in step S1, after the artificial intelligence-based article AI illustration method extracts the text summary and the keyword list, a preset blacklist and TD-IDF word frequency weight verification algorithm are used to perform blacklist filtering and TD-IDF word frequency weight verification on the extracted text summary and keyword list, thereby obtaining an accurate text summary and keyword list.

[0014] According to one embodiment of the artificial intelligence-based article AI illustration method of the present invention, after completing blacklist filtering and TD-IDF word frequency weight verification, the artificial intelligence-based article AI illustration method performs embedding text vectorization processing on the text summary and keyword list after blacklist filtering and TD-IDF word frequency weight verification, thereby obtaining text semantics for retrieving private image libraries.

[0015] According to an embodiment of the AI-powered article illustration method of the present invention, the AI-powered article illustration method, when using the large model of text-based images for intelligent illustration generation, includes the following steps:

[0016] Step C1: selecting a style of the image to be generated based on the text content of the image to be generated;

[0017] Step C2: Generate image generation prompt words based on the extracted text summary, keyword list, and image style;

[0018] Step C3: Input the generated picture generation prompt words into the text-image model for intelligent picture generation, and judge whether there is a suitable picture in the generated picture image list; if so, directly insert the picture image at the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture image list until a suitable picture is obtained.

[0019] According to an embodiment of the AI-powered article illustration method of the present invention, the AI-powered article illustration method uses a large image-based model to generate intelligent illustrations, including the following steps:

[0020] Step D1: Selecting a style of illustration to be generated based on the text content of the illustration to be generated;

[0021] Step D2: Generate a picture generation prompt word based on the extracted text summary, keyword list, picture style, and picture generation prompt word structure; wherein the picture generation prompt word structure includes: main body description + style / art style + detail adjustment + control parameters;

[0022] Step D3: Input the generated picture generation prompt words into the picture generation model for intelligent picture generation, and judge whether the generated picture list is directly used; if so, insert the picture directly into the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture list until a suitable picture is obtained.

[0023] According to an embodiment of the AI ​​picture matching method for articles based on artificial intelligence of the present invention, in step S3, after determining that there is no reference picture, the AI ​​picture matching method for articles based on artificial intelligence generates prompt words by modifying the picture matching, or regenerates the picture matching prompt words with reference to the currently generated picture matching picture list as a reference picture, and inputs the words into the raw picture model to regenerate the picture matching picture list until there is a suitable picture in the regenerated picture matching picture list.

[0024] According to an embodiment of the AI-based article picture matching method of the present invention, in step S4, after the AI-based article picture matching method generates a picture matching list, the picture matching list is reviewed to determine whether there are suitable pictures through the following steps:

[0025] Step F1: Use the CLIP model to encode the text and the pictures in the picture list to obtain text feature vectors and picture feature vectors;

[0026] Step F2: Use the cosine similarity in the kNN algorithm to measure the distance between the text feature vector and the image feature vector, and sort the images based on the distance similarity;

[0027] Step F3: Based on the sorted image list, select whether there is a suitable image. If so, directly insert the image into the corresponding article of the text content to be illustrated, and end the article intelligent illustration process; if not, continue to generate the image list until a suitable image is found and inserted into the corresponding position of the text content to be illustrated, and complete the article intelligent illustration process.

[0028] The present invention also provides a computer-readable medium storing computer program code, which implements the method described above when executed by a processor.

[0029] The present invention also provides an AI-based article illustration device, comprising:

[0030] a memory for storing instructions executable by the processor; and

[0031] A processor is configured to execute the instructions to implement the method described above.

[0032] Compared with the existing technology, the present invention has the following beneficial effects: in order to realize intelligent illustration of articles, the present invention adopts large-scale model deep learning technology to accurately extract text summaries and keyword lists, and then uses the extracted text summaries and keyword lists as natural semantics for image retrieval and intelligent image generation. Multiple rounds of image generation are performed through an advanced large-scale image generation model, thereby finally completing the intelligent illustration of articles. Compared with the existing technology, the present invention breaks through the limitations of traditional illustration and combines artificial intelligence with text semantic analysis. It not only solves the long-standing problems of low illustration efficiency, high labor costs, and copyright conflicts, but also makes the final selected images highly matched with the text content through deep semantic understanding, thereby improving the fit between the illustration and the article theme, thereby greatly enriching the reader's visual experience and greatly improving the overall quality of the content. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The above features and advantages of the present invention will be better understood after reading the detailed description of the embodiments of the present disclosure in conjunction with the following drawings. In the drawings, the components are not necessarily drawn to scale, and components with similar related properties or characteristics may have the same or similar reference numerals.

[0034] Figure 1 It is a flowchart showing the steps of an embodiment of the artificial intelligence-based article AI illustration method of the present invention.

[0035] Figure 2It is a flowchart showing the steps of an embodiment of the method for intelligently generating illustrations using a large model of cultural images of the present invention.

[0036] Figure 3 It is a flowchart showing the steps of an embodiment of the method for intelligently generating an image using a large image model of the present invention.

[0037] Figure 4 It is a flowchart showing the steps of an embodiment of the method for reviewing an accompanying picture list of the present invention. DETAILED DESCRIPTION

[0038] To more clearly illustrate the technical solutions of the embodiments of this application, the following is a brief introduction to the drawings required for describing the embodiments. Obviously, the drawings described below are merely examples or embodiments of this application. Those skilled in the art can apply this application to other similar scenarios based on these drawings without inventive effort. Unless otherwise apparent from the context or otherwise noted, the same reference numerals in the figures represent the same structure or operation.

[0039] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0040] Unless otherwise specifically stated, the relative arrangement of the parts and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present application. At the same time, it should be understood that, for ease of description, the sizes of the various parts shown in the drawings are not drawn according to actual proportional relationships. The techniques, methods and equipment known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, the techniques, methods and equipment should be considered as part of the authorization specification. In all examples shown and discussed here, any specific values ​​should be interpreted as being merely exemplary and not as limitations. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that similar numbers and letters represent similar items in the following figures, and therefore, once an item is defined in one figure, it does not need to be further discussed in subsequent figures.

[0041] When describing the embodiments of the present invention, for ease of explanation, cross-sectional views illustrating device structures may be partially enlarged and not to scale. Furthermore, these schematic views are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, three-dimensional dimensions, including length, width, and depth, should be included.

[0042] In the description of this application, it should be understood that the directions or positional relationships indicated by directional words such as "front, back, up, down, left, right", "horizontal, vertical, vertical, horizontal" and "top, bottom" are usually based on the directions or positional relationships shown in the accompanying drawings. They are only for the convenience of describing this application and simplifying the description. Unless otherwise specified, these directional words do not indicate or imply that the device or element referred to must have a specific direction or be constructed and operated in a specific direction. Therefore, they cannot be understood as limiting the scope of protection of this application; the directional words "inside and outside" refer to the inside and outside relative to the outline of each component itself.

[0043] An embodiment of an AI-based method for illustrating articles is disclosed herein. Figure 1 This is a flow chart showing an embodiment of the AI-based article illustration method of the present invention. Figure 1 , the following is a detailed description of the steps of the AI ​​illustration method for articles based on artificial intelligence.

[0044] Step S1: Select the text content to be matched with the image, and input the selected text content into a preset language extraction model to extract the text summary and keyword list.

[0045] In this embodiment, after the user selects the text content to be illustrated in the image editor, the selected text content is inserted into a preset prompt word template. Then, based on the prompt word template, prompt words are assembled to form the text summary and keyword list to be extracted. The assembled prompt words are then input into a language extraction large language model (such as Tongyi Qianwen, DeepSeek, etc.). Leveraging the powerful natural language processing capabilities of the language extraction large language model, the text summary and keyword list extracted by the language extraction large language model are obtained. The conditions for extracting keywords and text summaries are as follows:

[0046] 1. The generated text summary must be concise, between 2-4 sentences, and focus on the core information;

[0047] 2. The generated text summary contains a list of 3 to 10 keywords covering the main topic, key entities, and important details, arranged in descending order of importance;

[0048] 3. The generated text summary must remain objective and neutral, avoiding subjective interpretation;

[0049] 4. The generated keywords should be in Chinese (proper nouns can be retained in English), taking into account both the overall theme and specific details, and separated by commas;

[0050] 5. Prioritize important concepts, taking into account both the overall theme and specific details.

[0051] For example, the preset text content to be matched with an image is: "The plum blossoms complement the gray tiles, white walls, and winding corridors of the garden, creating a unique "Suzhou-style plum blossom viewing" culture. In Suzhou gardens, many pavilions and towers are named after "plum", adding a touch of elegance to the gardens. In the Humble Administrator's Garden, you can enjoy large areas of planted plum blossoms at the Xuexiang Yunwei Pavilion and on the west side of Zhuiyun Peak. The Xuexiang Yunwei Pavilion is a quaint square pavilion. "Xuexiang" refers to the fragrance of plum blossoms, and "Yunwei" means lush. The white plum blossoms here are as beautiful as snow and as beautiful as clouds and mist." After the corresponding prompt words are obtained by assembling the prompt word template and inputting them into the language extraction model, the language extraction model output is as follows:

[0052] Text Abstract: The "Suzhou-style plum blossom appreciation" culture in Suzhou gardens is unique, with pavilions and towers mostly named after "plum". Among them, the Xuexiang Yunwei Pavilion and the west side of Zhuiyun Peak in the Humble Administrator's Garden are planted with large tracts of plum blossoms, showcasing the simplicity and elegance of plum blossoms.

[0053] Keyword list: Suzhou gardens, Suzhou-style plum blossom viewing, Humble Administrator's Garden, Xuexiang Yunwei Pavilion, plum blossoms, elegance, ground-planted plum blossoms, white plum blossoms, simplicity.

[0054] In addition, in this embodiment, after the text summary and keyword list are extracted based on the text content to be matched with the image, a preset blacklist is used to filter out invalid and sensitive words from the extracted text summary and keyword list, and then the TD-IDF word frequency weight verification algorithm is used to verify the TD-IDF word frequency weight, so as to obtain an accurate text summary and keyword list.

[0055] Specifically, in this embodiment, after filtering invalid and sensitive words through a preset blacklist, the TF-IDF algorithm is used to verify the keyword weight distribution and remove low-frequency or invalid keywords. The TF-IDF algorithm uses a corpus compiled from its own historical material library as the corpus for the IDF Inverse Document Frequency index. This step is effective in removing low-frequency keywords, and the summarized keywords are more accurate.

[0056] In addition, in this embodiment, after completing the blacklist filtering and TD-IDF word frequency weight verification, the final text summary and keyword list are embedded with text vectorization to obtain the text semantics (i.e., text vector) used to retrieve the private gallery. Then, based on the text semantics, the private gallery is directly retrieved through the multimodal vector feature value, and it is determined whether to use the pictures in the private gallery.

[0057] Step S2: Determine whether to use a private gallery; if so, retrieve a list of accompanying pictures from the private gallery based on the extracted text summary and keyword list; if not, generate a list of accompanying pictures based on a preset raw picture model.

[0058] In this embodiment, after extracting the text summary and keyword list through the above steps, an editor manually reviews the private library in the image list retrieved from the private library and searches for matching images based on the text summary and keyword list. If no suitable images are found, a raw image model (such as the Tongyi Wanxiang or Doubao-Wen raw image model) is used to generate a matching image list based on the text summary and keyword list.

[0059] The large-scale raw image model includes a large-scale text-generated image model and a large-scale image-generated image model. After determining that a private image library is not currently in use, image generation prompts are generated based on the extracted text summary and keyword list. Then, intelligent image generation is performed using the large-scale text-generated image model or the large-scale image-generated image model based on the image generation prompts, resulting in a list of matching images. Figure 2 This is a flowchart showing the steps of an embodiment of the method for generating intelligent pictures using the large model of cultural images of the present invention. Figure 2 , describes in detail the steps of intelligent image generation using the Wenshengtu model.

[0060] Step C1: Select the style of the image to be generated based on the text content of the image to be generated.

[0061] Step C2: Generate image generation prompt words based on the extracted text summary, keyword list, and image style.

[0062] Step C3: Input the generated picture generation prompt words into the text-image model for intelligent picture generation, and judge whether there is a suitable picture in the generated picture image list; if so, directly insert the picture image at the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture image list until a suitable picture is obtained.

[0063] Specifically, in this embodiment, when the large model of text and image is used for intelligent image generation, the style of the image to be generated is first selected according to the text requirements, including photography, cartoon, ancient style, Chinese ink painting, simple drawing, science fiction and future style, picture book style, 3D rendering, etc. Then the text summary and keyword list generated by the language extraction large model are automatically filled into the prompt word template of the raw image to generate the corresponding image generation prompt word. Among them, the image generation prompt word is generated by two parts. One part is the preset picture style prompt word. For example, if the "photography" style is selected in this example, the system automatically adds the "photography" style prompt word in front of the "creative description": "architectural photography style, warm colors, low saturation". The other part is the image generation prompt word composed of the text summary and keyword list generated by the large model of language extraction.

[0064] In addition, in this embodiment, when using the picture style, text summary and keyword list to form the raw picture prompt, you can also modify or edit the raw picture prompt words as needed. For example, for a more specific detailed description, add "some trees and blue sky can be seen in the background, the sun shines on the building, and the glass curtain wall reflects the surrounding scenery and light." and other descriptions. Finally, after completing the editing of the raw picture prompt words, click "Start Creation" to generate intelligent picture matching. The text-based picture model will generate a corresponding list of generated picture matching images based on the picture matching prompt words, and you can directly click "Select" to insert the generated picture matching into the article to complete the intelligent picture matching. Among them, in this link, you can repeatedly modify the prompt words to generate the picture matching until you get a satisfactory picture matching.

[0065] Figure 3 This is a flowchart showing the steps of an embodiment of the method for generating intelligent pictures using a large picture model of the present invention. Figure 3 ,describe in detail, the detailed steps of intelligent image generation by the large-scale image model.

[0066] Step D1: Select the style of the image to be generated based on the text content of the image to be generated.

[0067] Step D2: Generate picture generation prompt words based on the extracted text summary, keyword list, picture generation style, and picture generation prompt word structure; wherein the picture generation prompt word structure includes: main body description + style / painting style + detail adjustment + control parameters.

[0068] Step D3: Input the generated picture generation prompt words into the picture generation model for intelligent picture generation, and judge whether the generated picture list is directly used; if so, insert the picture directly into the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture list until a suitable picture is obtained.

[0069] Specifically, in this embodiment, when constructing the image generation model to generate prompt words, the requirements are as follows:

[0070] 1. Description of the subject:

[0071] 1.1: Core content: Identify the core elements of the reference image (such as people, objects, and scenes) and indicate whether they need to be retained or modified.

[0072] 1.2: Keyword examples:

[0073] Preserved part: Based on the reference image, keep the original character pose and background structure;

[0074] Modifications: Changed the character's clothing to a cyberpunk style and replaced the forest with a futuristic city.

[0075] 2. Style / art style;

[0076] 2.1: Core content: Specify the art style or visual effect you want to convert.

[0077] 2.2: Keyword examples:

[0078] Common styles: anime style, photorealism, cyberpunk, watercolor;

[0079] Artist Style: by Studio Ghibli, in the style of Van Gogh, ArtStationtrending;

[0080] Lighting and tones: dramatic lighting, warm colors, neon lights.

[0081] 3.Detail adjustments:

[0082] 3.1: Core content: Fine-tune the picture details (such as texture, dynamics, atmosphere, etc.).

[0083] 3.2: Keyword examples:

[0084] Materials: Complex mechanical details, translucent silk fabric;

[0085] Dynamic: dynamic posture, flying petals, smoke effect;

[0086] Atmosphere: Mysterious fog, dreamy atmosphere.

[0087] 4. Parameter control (optional):

[0088] 4.1: Core content: Adjust the generation strength or similarity with the reference image through parameters.

[0089] 4.2: Common parameters (the syntax may be different for different models):

[0090] Denoising strength: 0.6 (the lower the value, the closer it is to the original image, 0.3-0.7 is commonly used);

[0091] CFG scale: 7 (controls the weight of prompt words, 7-12 is commonly used);

[0092] Seed: 1234 (fixed seed for reproducible results).

[0093] Step S3: Based on the generated illustration picture list, determine whether to use the picture directly; if so, directly insert the illustration picture at the corresponding position of the text content to be illustrated to complete the article intelligent illustration process; if not, determine whether there is a reference picture, if so, regenerate the illustration picture list based on the reference picture using the raw picture model; if not, return to step S2 to generate the illustration picture list based on the preset raw picture model.

[0094] In this embodiment, after the picture list is generated through the above steps, when it is determined that there is no reference picture, the picture generation prompt word is modified, or the picture generation prompt word is regenerated with reference to the currently generated picture list as a reference picture, and is input into the raw picture model to regenerate the picture list, until there is a suitable picture in the regenerated picture list.

[0095] Step S4: Based on the regenerated picture list, determine whether there is a suitable picture; if so, directly insert the picture into the corresponding article of the text content to be illustrated, and end the article intelligent illustration process; if not, continue to generate the picture list based on the regenerated picture list until there is a suitable picture and insert it into the corresponding position of the text content to be illustrated, and complete the article intelligent illustration process.

[0096] In this embodiment, after the picture list is regenerated through the above steps, the picture list needs to be reviewed to determine whether there are suitable pictures. Figure 4 This is a flowchart showing the steps of an embodiment of the method for reviewing a list of accompanying pictures of the present invention. Figure 4 , which details the steps for reviewing the image list.

[0097] Step F1: Use the CLIP model to encode the text and the pictures in the picture list to obtain text feature vectors and picture feature vectors.

[0098] Step F2: Use the cosine similarity in the kNN algorithm to measure the distance between the text feature vector and the image feature vector, and sort the images in sequence based on the distance similarity.

[0099] Step F3: Based on the sorted image list, select whether there is a suitable image. If so, directly insert the image into the corresponding article of the text content to be illustrated, and end the article intelligent illustration process; if not, continue to generate the image list until a suitable image is found and inserted into the corresponding position of the text content to be illustrated, and complete the article intelligent illustration process.

[0100] Specifically, in this embodiment, after using a large raw image model to generate multiple images at once to obtain a list of accompanying images, for example, generating five images each time, after the five images are generated, the CLIP (Contrastive Language-Image Pre-training) model is used to encode the text and generated images to obtain text feature vectors and image feature vectors. The text feature vector is then used together with the five generated image feature vectors to perform a distance measurement based on the cosine similarity in the kNN algorithm, and the images are sorted in descending order of similarity. Finally, based on the high-to-low semantic similarity between the images and the text as described above, a suitable accompanying image is manually selected in combination with the actual accompanying image intent. If no satisfactory accompanying image is found, the raw image prompt word is modified again to generate an accompanying image, and this cycle is repeated until a satisfactory accompanying image is generated for the article.

[0101] This specification also provides a computer-readable medium storing a computer program code, which, when executed by a processor, implements the above-mentioned artificial intelligence-based article AI illustration method.

[0102] This specification also provides an AI-based article illustration device based on artificial intelligence, including an instruction memory for storing instructions executable by a processor, and a processor for executing instructions in the instruction memory to implement the AI-based article illustration method based on artificial intelligence as described above.

[0103] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0104] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithmic steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. A skilled person may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.

[0105] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or executed using a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0106] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor so that the processor can read and write information from / to the storage medium. In an alternative, the storage medium may be integrated into the processor. The processor and storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and storage medium may reside in a user terminal as discrete components.

[0107] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software as a computer program product, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one location to another. A storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, where disks typically reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

Claims

1. An AI-based article illustration method, characterized in that: The following steps are involved: Step S1: Select the text content to be matched with the image, and input the selected text content into the preset language extraction model to extract the text summary and keyword list. The extracted text summary and keyword list are blacklist filtered and TD-IDF word frequency weight verification is performed using the preset blacklist and TD-IDF word frequency weight verification algorithm; Step S2: Determine whether to use a private image library; if so, retrieve a list of accompanying images from the private image library based on the extracted text summary and keyword list; if not, generate a list of accompanying images based on a preset raw image model; Step S3: Based on the generated picture list, determine whether to directly use the picture; if so, directly insert the picture at the corresponding position of the text content to be illustrated to complete the article intelligent picture matching process; if not, determine whether there is a reference picture, and if so, regenerate the picture list based on the reference picture using the raw picture model; if not, return to step S2 to generate the picture list based on the preset raw picture model; Step S4: Based on the regenerated picture list, determine whether there is a suitable picture; if so, directly insert the picture into the corresponding article of the text content to be illustrated, and end the article intelligent picture matching process; if not, continue to generate a picture list based on the regenerated picture list until a suitable picture is found and inserted into the corresponding position of the text content to be illustrated, and complete the article intelligent picture matching process; Multiple rounds of image generation are performed, and the CLIP model is used to encode text and images in the generated image list. The text and images are automatically sorted based on cosine similarity, and the image that best matches the text is selected and inserted into the corresponding position. Among them, in step S2, the raw image large model includes a text-generated image large model and a picture-generated image large model; wherein, the artificial intelligence-based article AI illustration method determines that the private image library is not currently in use, generates illustration generation prompt words based on the extracted text summary and keyword list, and then uses the text-generated image large model or the picture-generated image large model to perform intelligent illustration generation based on the illustration generation prompt words, thereby obtaining an illustration picture list.

2. The AI-based article illustration method according to claim 1 is characterized in that: In step S1, after the artificial intelligence-based article AI illustration method extracts the text summary and keyword list, it uses a preset blacklist and TD-IDF word frequency weight verification algorithm to perform blacklist filtering and TD-IDF word frequency weight verification on the extracted text summary and keyword list, thereby obtaining an accurate text summary and keyword list.

3. The AI-based article illustration method according to claim 2 is characterized in that: After completing blacklist filtering and TD-IDF word frequency weight verification, the artificial intelligence-based article AI illustration method performs embedding text vectorization processing on the text summary and keyword list after blacklist filtering and TD-IDF word frequency weight verification, thereby obtaining text semantics for retrieving private image libraries.

4. The AI-based article illustration method according to claim 1 is characterized in that: The AI-based article illustration method based on artificial intelligence uses the large-scale article-image model to generate intelligent illustrations, including the following steps: Step C1: selecting a style of the image to be generated based on the text content of the image to be generated; Step C2: Generate image generation prompt words based on the extracted text summary, keyword list, and image style; Step C3: Input the generated picture generation prompt words into the text-image model for intelligent picture generation, and judge whether there is a suitable picture in the generated picture image list; if so, directly insert the picture image at the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture image list until a suitable picture is obtained.

5. The AI-based article illustration method according to claim 1 is characterized in that: The AI-based article illustration method based on artificial intelligence uses a large-scale image generation model to generate intelligent illustrations, including the following steps: Step D1: Selecting a style of illustration to be generated based on the text content of the illustration to be generated; Step D2: Generate a picture generation prompt word based on the extracted text summary, keyword list, picture style, and picture generation prompt word structure; wherein the picture generation prompt word structure includes: main body description + style / art style + detail adjustment + control parameters; Step D3: Input the generated picture generation prompt words into the picture generation model for intelligent picture generation, and judge whether the generated picture list is directly used; if so, insert the picture directly into the corresponding position of the text content to be illustrated to complete the article intelligent picture generation process; if not, modify the picture generation prompt words and continue to generate the picture list until a suitable picture is obtained.

6. The AI-based article illustration method according to claim 1, characterized in that: In step S3, after the artificial intelligence-based article AI illustration method determines that there is no reference picture, it generates prompt words by modifying the illustration, or regenerates the illustration generation prompt words with reference to the currently generated illustration picture list as a reference picture, and inputs them into the raw picture model to regenerate the illustration picture list until there is a suitable picture in the regenerated illustration picture list.

7. The AI-based article illustration method according to claim 1, characterized in that: In step S4, after the AI-based article picture matching method generates a picture matching list, the picture matching list is reviewed to determine whether there are suitable pictures through the following steps: Step F1: Use the CLIP model to encode the text and the pictures in the picture list to obtain text feature vectors and picture feature vectors; Step F2: Use the cosine similarity in the kNN algorithm to measure the distance between the text feature vector and the image feature vector, and sort the images based on the distance similarity; Step F3: Based on the sorted image list, select whether there is a suitable image. If so, directly insert the image into the corresponding article of the text content to be illustrated, and end the article intelligent illustration process; if not, continue to generate the image list until a suitable image is found and inserted into the corresponding position of the text content to be illustrated, and complete the article intelligent illustration process.

8. A computer-readable medium storing computer program code, characterized in that: When the computer program code is executed by a processor, the computer program code implements the method according to any one of claims 1 to 7.

9. An AI-based article illustration device, characterized in that: include: a memory for storing instructions executable by the processor; as well as A processor, configured to execute the instructions to implement the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Image-text content generation method and device, equipment and storage medium

    CN117032869A