A large model-based image-text work generation method and device

By using large-scale model-based image and text generation technology, the content and illustrations for handwritten newspapers are automatically generated, solving the problem of time-consuming and labor-intensive production of handwritten newspapers, improving efficiency and quality, and enhancing students' creative experience.

CN122220591APending Publication Date: 2026-06-16北京爱宾果科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京爱宾果科技有限公司
Filing Date
2026-05-18
Publication Date
2026-06-16

AI Technical Summary

Technical Problem

Creating student-made posters is time-consuming and laborious, and due to limited drawing or layout skills, the quality varies, affecting learning outcomes and teacher evaluations.

Method used

Employing a large-model-based image and text generation technology, it automatically generates text content and accompanying images, intelligently typesets them, simulates user handwriting, and generates image and text works that conform to paper properties.

Benefits of technology

It significantly improves the efficiency of creating handwritten newspapers, retains the creative participation, generates high-quality graphic works, and reduces the workload of students.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122220591A_ABST
    Figure CN122220591A_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data processing, and particularly relates to a picture-text work generation method and device based on a large model. The method comprises the following steps: S1, obtaining a picture-text work outline given by a user; S2, generating a plurality of picture-text materials matched with the composition outline based on a pre-trained content generation large model, each picture-text material comprising a plurality of text materials and picture materials matched with the text materials; S3, performing layout on each picture-text material according to a randomly given layout style to generate a picture-text work; and S4, displaying the multiple picture-text works after layout, so as to be modified and selected by the user. The application can automatically generate text and pictures, greatly shorten the picture-text work production time, help students quickly complete homework, and meanwhile retain a certain degree of creative participation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of data processing technology, and specifically relates to a method and apparatus for generating graphic works based on a large model. Background Technology

[0002] In the current educational environment, students generally face a heavy academic burden, among which handwritten newspaper assignments, due to their comprehensive nature and long processing time, have become a significant source of this burden for many students. Traditional handwritten newspaper creation requires students to manually write text, collect and draw pictures, and design the layout. The entire process is time-consuming and laborious, especially as the deadline approaches, students often struggle to complete high-quality work, consuming a large amount of their rest time with unsatisfactory results. Furthermore, some students' limited drawing or layout skills lead to inconsistent quality in their handwritten newspapers, affecting learning outcomes and teacher evaluations. Summary of the Invention

[0003] To address the aforementioned issues, this application provides a method and apparatus for generating graphic works based on large models. The graphic generation technology based on large models assists students in quickly generating text content and accompanying images, and intelligently typeset the layout, significantly improving the efficiency of creating handwritten newspapers.

[0004] The first aspect of this application provides a method for generating graphic works based on a large model, mainly including:

[0005] Step S1: Obtain the outline of the graphic and textual work provided by the user;

[0006] Step S2: Based on the pre-trained content generation model, generate multiple text and image materials that match the essay outline. Each text and image material includes several text materials and image materials that match the text materials.

[0007] Step S3: Arrange each piece of text and image material according to the randomly given layout style to generate a text and image work;

[0008] Step S4: Display the multiple layoutd graphic works for users to modify and select;

[0009] Step S2 further includes:

[0010] Step S21: Calculate the offset parameters of the handwritten characters relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user.

[0011] Step S22: Generate an artistic font suitable for the user based on the bias parameters;

[0012] Step S23: Optimize the text material using the selected artistic font.

[0013] Preferably, step S1 further includes:

[0014] Step S11: Obtain the outline of the graphic and textual works displayed by the user in the scanning area using a scanning device;

[0015] Step S12: Identify the themes and keywords in the outline of the graphic works.

[0016] Preferably, the offset parameters include a first offset parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second offset parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

[0017] Preferably, step S3 further includes:

[0018] Step S31: Obtain the paper attributes of the graphic works set by the user, including paper size and number of columns;

[0019] Step S32: Combine, split, or optimize the multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can adapt to the requirements of paper size and number of columns.

[0020] Preferably, step S4 further includes:

[0021] Step S41: Display the typed graphic works in the projection area one by one;

[0022] Step S42: In response to the user's selection, movement, or deletion of a specified section in the displayed graphic and text works, the graphic and text works are rearranged and displayed.

[0023] Preferably, step S4 further includes:

[0024] Step S43: Display at least two typeset graphic works simultaneously in the projection area;

[0025] Step S44: In response to the user's operation of exchanging columns within different graphic works, the two columns belonging to different graphic works are swapped.

[0026] The second aspect of this application provides a graphic and textual work generation device based on a large model, mainly comprising:

[0027] The "Image and Text Work Outline Acquisition Module" is used to acquire the image and text work outline provided by the user.

[0028] The material generation module is used to generate multiple text and image materials that match the essay outline based on a pre-trained content generation model. Each text and image material includes several text materials and image materials that match the text materials.

[0029] The graphic and text creation module is used to format each piece of graphic and text material according to a randomly given layout style and generate graphic and text works.

[0030] The editing module displays multiple formatted graphic and text works for users to modify and select.

[0031] The material generation module includes:

[0032] The bias parameter acquisition unit is used to calculate the bias parameters of the handwritten text strokes relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user.

[0033] An artistic font generation unit is used to generate an artistic font suitable for the user based on the bias parameters.

[0034] An optimization processing unit is used to optimize the text material using a selected artistic font.

[0035] Preferably, the graphic and textual work outline acquisition module includes:

[0036] The sketch acquisition unit is used to acquire the outline of the graphic and textual works displayed by the user in the scanning area through the scanning device;

[0037] The theme and keyword recognition unit is used to identify the themes and keywords in the outline of graphic works.

[0038] Preferably, the offset parameters include a first offset parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second offset parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

[0039] Preferably, the graphic and text creation module includes:

[0040] The style generation unit is used to obtain the paper attributes of the graphic works set by the user, including paper size and number of columns.

[0041] The typesetting unit is used to combine, split, or optimize multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can fit the requirements of paper size and number of columns.

[0042] Preferably, the correction module includes:

[0043] The single-work display unit is used to sequentially display the laid-out graphic and text works in the projection area;

[0044] The content correction unit is used to respond to the user's selection, movement, or deletion of a specified column in the displayed graphic and text works, and to rearrange and display the graphic and text works.

[0045] Preferably, the correction module includes:

[0046] The multi-work display unit is used to simultaneously display at least two typed graphic works in the projection area;

[0047] The multi-work content adjustment unit is used to respond to users' operations on exchanging columns within different graphic and text works, allowing two columns belonging to different graphic and text works to be swapped.

[0048] This application can automatically generate text and images, significantly shortening the production time of graphic works and helping students complete assignments quickly while retaining a certain degree of creative participation. Attached Figure Description

[0049] Figure 1 This is a flowchart of a preferred implementation of the graphic and textual work generation method based on a large model in this application.

[0050] Figure 2 This is a schematic diagram showing the scanned area.

[0051] Figure 3 This is a schematic diagram of paper properties design. Detailed Implementation

[0052] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The described embodiments are only some, not all, of the embodiments of this application. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0053] According to the first aspect of this application, a method for generating graphic works based on a large model, such as... Figure 1 As shown, it mainly includes:

[0054] Step S1: Obtain the outline of the graphic and textual works provided by the user.

[0055] This step is used to obtain the outline of the graphic work, mainly referring to hand-drawn sketches for handwritten newspapers, posters, etc. The methods of obtaining this outline include, but are not limited to, manual input, photo recognition, or scanning. This application prefers to use scanning. After projecting the learning machine content onto the desktop, a scanning area is provided in addition to the control buttons, such as... Figure 2 As shown, users can place their hand-drawn sketches, posters, or other hand-drawn artwork on the scanning area and then start the scanning program to obtain the outline of the artwork.

[0056] Users of this application include students, teachers, or parents.

[0057] In some alternative implementations, step S1 further includes:

[0058] Step S11: Obtain the outline of the graphic and textual works displayed by the user in the scanning area using a scanning device;

[0059] Step S12: Identify the themes and keywords in the outline of the graphic works.

[0060] This embodiment provides a specific scanning method for obtaining an outline of a graphic work. The recognition model, based on a pre-trained handwritten Chinese character recognition model, can easily and quickly identify the text within the scanned image. For example, the offline recognition process includes three stages: preprocessing, feature extraction, and classification. Online recognition uses a Hidden Markov Model to handle stroke connections and stroke order variations, and improves robustness through dynamic time warping. For hand-drawn object graphics, a target object recognition model can be used for rapid recognition, and text keywords identified by the character recognition model can be superimposed as auxiliary features to improve the classification accuracy of the object graphics, obtain object graphic keywords, and comprehensively form the theme and keywords of the graphic work outline.

[0061] Step S2: Based on the pre-trained content generation model, generate multiple text and image materials that match the essay outline. Each text and image material includes several text materials and image materials that match the text materials.

[0062] After the system reads the outline of the graphic and textual work, it can generate multiple graphic and textual materials based on the theme and keywords using a content generation model. For a handwritten graphic and textual work on a reflection on Journey to the West, the generated graphic and textual materials can include a brief introduction to the work, an introduction to the author, a brief introduction to the characters' relationships, a brief introduction to the plot, etc.

[0063] In some alternative implementations, step S2 further includes:

[0064] Step S21: Calculate the offset parameters of the handwritten characters relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user.

[0065] Step S22: Generate an artistic font suitable for the user based on the bias parameters;

[0066] Step S23: Optimize the text material using the selected artistic font.

[0067] In this embodiment, by changing the conventional printed font to an artistic font, graphic works such as handwritten newspapers and posters can better reflect the actual hand-drawn effect, while also enhancing their aesthetics. The artistic font can use existing artistic font templates, or it can be designed by the user and pre-stored in the system font library.

[0068] Each user of this application can set an artistic font suitable for showcasing their own style, or, as described in this embodiment, construct an artistic font similar to their own handwritten font. Specifically, the glyph characteristics of each font within a user-inputted handwritten graphic artwork outline can be determined, and the glyphs of the standard font can be modified based on these characteristics to form the aforementioned artistic font.

[0069] In some optional implementations, the bias parameters include a first bias parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second bias parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

[0070] In this embodiment, based on the above scaling parameters and the structural composition of the characters, the position of each stroke in each character of the artistic font can be determined. Then, based on the first offset parameter, the center of gravity of each stroke is adjusted. Finally, based on the second offset parameter, the tilt angle and bending parameters of the strokes are adjusted.

[0071] It should be noted that the first bias parameter and the scaling parameter are both single parameters, while the second bias parameter is a combination of multiple parameters. The more second bias parameters used, the smoother the strokes will be. Alternatively, a small number of second bias parameters can be used, and the strokes can be smoothed using spline interpolation.

[0072] Understandably, the above methods can simulate user handwriting, and further enable the use of industrial control software to control writing and automatically generate hand-drawn graphic works.

[0073] Step S3: Arrange each piece of text and image material according to the randomly given layout style to generate a text and image work.

[0074] After providing multiple text and image materials, the layout can be arranged according to the preset layout style. In this step, users can design the layout style in the background, such as selecting the paper size, layout selection, word limit, and column number settings.

[0075] In some alternative implementations, step S3 further includes:

[0076] Step S31: Obtain the paper attributes of the graphic works set by the user, including paper size and number of columns;

[0077] Step S32: Combine, split, or optimize the multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can adapt to the requirements of paper size and number of columns.

[0078] The paper type used in this application is, for example, A4 paper, landscape orientation, and the number of columns is, for example, three rows per column, or four rows per two columns, etc. Figure 3 As shown, typically, each piece of graphic and textual material may generate a lot of content based on the keywords in the outline of the graphic and textual work. This requires further deletion or combination in step S32. The system selects several pieces of text or image materials with high suitability based on the suitability of the keywords and the suitability of the theme, and finally forms multiple graphic and textual works.

[0079] Finally, in step S4, the multiple layoutd graphic works are displayed for users to modify and select.

[0080] It should be noted that the system can select several high-quality images and texts from the generated works for display. The quality score Q of the images and texts can be described by the following formula:

[0081] Q = αR + βD + γC.

[0082] Where α, β, and γ are weighting coefficients, and α + β + γ = 1.

[0083] R represents the image-text relevance score, indicating the semantic matching degree between the image / text materials and the outline. It can be calculated using the cosine similarity of semantic vectors: R = (Vm·Vd) / (||Vm||*||Vd||), where Vm is the semantic vector of the outline text and Vd is the semantic vector of the image / text materials. The semantic vectors can be extracted using pre-trained language models, such as BERT or CLIP.

[0084] D stands for Diversity Score, representing the creative diversity of the text and image materials. It is calculated by counting the proportion of repeated keywords in multiple text and image materials. The lower the repetition rate, the higher the D score. For example, if 5 materials are generated and the keyword repetition rate is 20%, then D = 1 - 0.2 = 0.8.

[0085] C represents the coherence score, indicating the logical connection between the textual and visual materials. It is calculated using cross-modal matching scores, as shown in the following formula:

[0086] ;

[0087] in, The total number of paragraphs in the text material. Indicates the first The text of each paragraph With each picture The highest similarity value, specifically, the text and various pictures Input the text vector and image vector into a multimodal model, such as CLIP, and calculate their cosine similarity. For each paragraph, select the image that best matches it and score it.

[0088] Step S4 further includes: allowing users to make simple modifications to each graphic and textual work, such as adjusting the overall layout and the position of each text or image.

[0089] In some alternative implementations, step S4 further includes:

[0090] Step S41: Display the typed graphic works in the projection area one by one;

[0091] Step S42: In response to the user's selection, movement, or deletion of a specified section in the displayed graphic and text works, the graphic and text works are rearranged and displayed.

[0092] This example demonstrates modifications to a single graphic work, primarily involving deleting content from a section and adjusting the position of that content.

[0093] In some alternative implementations, step S4 further includes:

[0094] Step S43: Display at least two typeset graphic works simultaneously in the projection area;

[0095] Step S44: In response to the user's operation of exchanging columns within different graphic works, the two columns belonging to different graphic works are swapped.

[0096] This embodiment demonstrates how to modify multiple graphic and textual works simultaneously, mainly by swapping the sections in two graphic and textual works to improve the user experience.

[0097] The second aspect of this application provides a large-model-based graphic and textual work generation device corresponding to the above-mentioned method, mainly comprising:

[0098] The "Image and Text Work Outline Acquisition Module" is used to acquire the image and text work outline provided by the user.

[0099] The material generation module is used to generate multiple text and image materials that match the essay outline based on a pre-trained content generation model. Each text and image material includes several text materials and image materials that match the text materials.

[0100] The graphic and text creation module is used to format each piece of graphic and text material according to a randomly given layout style and generate graphic and text works.

[0101] The editing module displays multiple formatted graphic and text works for users to modify and select.

[0102] The material generation module includes:

[0103] The bias parameter acquisition unit is used to calculate the bias parameters of the handwritten text strokes relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user.

[0104] An artistic font generation unit is used to generate an artistic font suitable for the user based on the bias parameters.

[0105] An optimization processing unit is used to optimize the text material using a selected artistic font.

[0106] In some optional implementations, the graphic and textual work outline acquisition module includes:

[0107] The sketch acquisition unit is used to acquire the outline of the graphic and textual works displayed by the user in the scanning area through the scanning device;

[0108] The theme and keyword recognition unit is used to identify the themes and keywords in the outline of graphic works.

[0109] In some optional implementations, the bias parameters include a first bias parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second bias parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

[0110] In some optional implementations, the graphic and text creation module includes:

[0111] The style generation unit is used to obtain the paper attributes of the graphic works set by the user, including paper size and number of columns.

[0112] The typesetting unit is used to combine, split, or optimize multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can fit the requirements of paper size and number of columns.

[0113] In some alternative implementations, the correction module includes:

[0114] The single-work display unit is used to sequentially display the laid-out graphic and text works in the projection area;

[0115] The content correction unit is used to respond to the user's selection, movement, or deletion of a specified column in the displayed graphic and text works, and to rearrange and display the graphic and text works.

[0116] In some alternative implementations, the correction module includes:

[0117] The multi-work display unit is used to simultaneously display at least two typed graphic works in the projection area;

[0118] The multi-work content adjustment unit is used to respond to users' operations on exchanging columns within different graphic and text works, allowing two columns belonging to different graphic and text works to be swapped.

[0119] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for generating graphic and textual works based on a large model, characterized in that, include: Step S1: Obtain the outline of the graphic and textual work provided by the user; Step S2: Based on the pre-trained content generation model, generate multiple text and image materials that match the essay outline. Each text and image material includes several text materials and image materials that match the text materials. Step S3: Arrange each piece of text and image material according to the randomly given layout style to generate a text and image work; Step S4: Display the multiple layoutd graphic works for users to modify and select; Step S2 further includes: Step S21: Calculate the offset parameters of the handwritten characters relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user. Step S22: Generate an artistic font suitable for the user based on the bias parameters; Step S23: Optimize the text material using the selected artistic font.

2. The method for generating graphic and textual works based on a large model as described in claim 1, characterized in that, Step S1 further includes: Step S11: Obtain the outline of the graphic and textual works displayed by the user in the scanning area using a scanning device; Step S12: Identify the themes and keywords in the outline of the graphic works.

3. The method for generating graphic and textual works based on a large model as described in claim 1, characterized in that, The offset parameters include a first offset parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second offset parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

4. The method for generating graphic and textual works based on a large model as described in claim 1, characterized in that, Step S3 further includes: Step S31: Obtain the paper attributes of the graphic works set by the user, including paper size and number of columns; Step S32: Combine, split, or optimize the multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can adapt to the requirements of paper size and number of columns.

5. The method for generating graphic and textual works based on a large model as described in claim 1, characterized in that, Step S4 further includes: Step S41: Display the typed graphic works in the projection area one by one; Step S42: In response to the user's selection, movement, or deletion of a specified section in the displayed graphic and text works, the graphic and text works are rearranged and displayed.

6. The method for generating graphic and textual works based on a large model as described in claim 1, characterized in that, Step S4 further includes: Step S43: Display at least two typeset graphic works simultaneously in the projection area; Step S44: In response to the user's operation of exchanging columns within different graphic works, the two columns belonging to different graphic works are swapped.

7. A device for generating graphic and textual works based on a large model, characterized in that, include: The "Image and Text Work Outline Acquisition Module" is used to acquire the image and text work outline provided by the user. The material generation module is used to generate multiple text and image materials that match the essay outline based on a pre-trained content generation model. Each text and image material includes several text materials and image materials that match the text materials. The graphic and text creation module is used to format each piece of graphic and text material according to a randomly given layout style and generate graphic and text works. The editing module displays multiple formatted graphic and text works for users to modify and select. The material generation module includes: The bias parameter acquisition unit is used to calculate the bias parameters of the handwritten text strokes relative to the standard font strokes based on the hand-drawn graphic artwork outline input by the user. An artistic font generation unit is used to generate an artistic font suitable for the user based on the bias parameters. An optimization processing unit is used to optimize the text material using a selected artistic font.

8. The graphic and textual work generation device based on a large model as described in claim 7, characterized in that, The module for obtaining the outline of the graphic and textual works includes: The sketch acquisition unit is used to acquire the outline of the graphic and textual works displayed by the user in the scanning area through the scanning device; The theme and keyword recognition unit is used to identify the themes and keywords in the outline of graphic works.

9. The graphic and textual work generation device based on a large model as described in claim 7, characterized in that, The offset parameters include a first offset parameter of the center of gravity of the handwritten character stroke relative to the center of gravity of the standard font stroke, a second offset parameter of each set point on the handwritten character stroke relative to the center of gravity of the handwritten character stroke, and a scaling parameter of the distance between the start and end points of the handwritten character stroke relative to the distance between the start and end points of the standard font stroke.

10. The graphic and textual work generation device based on a large model as described in claim 7, characterized in that, The graphic and text creation module includes: The style generation unit is used to obtain the paper attributes of the graphic works set by the user, including paper size and number of columns. The typesetting unit is used to combine, split, or optimize multiple text materials involved in each graphic document so that the combined, split, or optimized text materials can fit the requirements of paper size and number of columns.

11. The graphic and textual work generation device based on a large model as described in claim 7, characterized in that, The correction module includes: The single-work display unit is used to sequentially display the laid-out graphic and text works in the projection area; The content correction unit is used to respond to the user's selection, movement, or deletion of a specified column in the displayed graphic and text works, and to rearrange and display the graphic and text works.

12. The graphic and textual work generation device based on a large model as described in claim 7, characterized in that, The correction module includes: The multi-work display unit is used to simultaneously display at least two typed graphic works in the projection area; The multi-work content adjustment unit is used to respond to users' operations on exchanging columns within different graphic and text works, allowing two columns belonging to different graphic and text works to be swapped.