AI-based cross-modal cartoon content generation and personalized recommendation method and system
By leveraging an AI engine to achieve cross-modal comic content generation and personalized recommendations in multi-person team collaboration, the challenge of comic content generation in multi-person team collaboration has been solved, improving creation efficiency and quality, and realizing personalized recommendations and creative richness.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2026-03-10
AI Technical Summary
Existing technologies struggle to generate and personalize comic content in collaborative multi-person teams, lacking effective technical solutions.
By using an AI-based cross-modal comic content generation and personalized recommendation method, and leveraging multi-user collaborative user groups and an AI engine, we can generate and adjust different modalities of content, including text, voice, and images, to collaboratively create comic content.
It improves the efficiency and quality of comic creation, enables personalized recommendations, enhances the flexibility and richness of creation, and ensures efficient collaboration in the creation process and multi-dimensional content protection.
Smart Images

Figure CN120823278B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of cartoon generation and generative interactive AI, and particularly relates to a cross-modal cartoon content generation and personalized recommendation method and system based on AI. BACKGROUND
[0002] The generative interactive AI is trained by a large amount of data and can understand and generate natural language text. When interacting with users, it can analyze user input and generate appropriate replies or perform specific tasks according to the context and pre-training knowledge, such as generating cartoons.
[0003] In related technologies, a multi-layer neural network is mainly constructed by deep learning and generative adversarial network (GAN) technology to extract features from a large-scale data set and learn, so that the AI can identify different elements of animation style, such as character design, background construction, color use, etc. GAN consists of a generator and a discriminator, the generator generates images, and the discriminator judges the authenticity of the images, and the two constantly compete, so that the generator generates more realistic animation pictures, making the cartoon generation more interactive and flexible. For example, users can gradually improve the theme, plot, character, and other elements of the cartoon through multiple rounds of dialogue with the AI. Related prior art includes a cartoon automatic generation method based on a material engine proposed by Chinese invention patent CN102110304B, an interactive cartoon generation system based on generative AI technology proposed by CN117391923A, and a cartoon image generation method proposed by CN119810223A.
[0004] Related cartoon projects of professional cartoon companies are usually completed by a team, a group, or even the entire company in collective collaboration. Each member in the team has different cartoon expertise, such as some people are good at text description, some people are good at cartoon drafts, and some people are good at music, etc. How to efficiently realize cartoon content generation and personalized recommendation under multi-person team collaboration, related technologies do not provide effective technical solutions. SUMMARY
[0005] To solve the above technical problems, the present application proposes a cross-modal cartoon content generation and personalized recommendation method and system based on AI.
[0006] In the first aspect of the present application, a cross-modal cartoon content generation method based on AI is proposed, which is applied to a multi-person collaborative user group, and the multi-person collaborative user group includes at least one first user and at least one second user.
[0007] The method comprises:
[0008] generating first control information by the first user, the first control information indicating that a first AI engine generates first cartoon content based on first modal content loaded by the first control information, and indicating that at least one second user responds to the first control information;
[0009] The first AI engine obtains response information of the second user to the first control information, and adjusts the first cartoon content based on the response information to generate second cartoon content.
[0010] The response information includes second modal content.
[0011] The first modal content and the second modal content include any one of the following modalities or a combination thereof: text, voice, and image.
[0012] In one scenario, the at least one second user is specified by the first user based on the first control information.
[0013] In another scenario, the at least one second user is automatically determined by the first AI engine based on the first modal content.
[0014] In a second aspect of the present application, an AI-based cross-modal cartoon content generation method is provided, which is applied to an interactive group containing N users {P1, P2, PN}, N>2, and the method includes the following steps:
[0015] When a user Pi generates a control message Ci, the control message Ci is sent to an AI engine Ai, and the engine Ai generates cartoon content Mi based on modal content Ti loaded by the control message Ci.
[0016] An interactive user group InterG is determined based on the control message Ci or the modal content Ti; the interactive user group InterG at least contains a user Pj and a user Pk; i, j, k∈{1, 2, 3, …, N} and i≠j≠k.
[0017] The engine Ai receives response information of the user Pj and the user Pk to the control message Ci, and adjusts the cartoon content Mi based on the response information to generate cartoon content M0.
[0018] Each of the N users {P1, P2, PN} corresponds to an AI engine;
[0019] The response message of the user Pa to the control message Ci is generated by the AI engine Aa corresponding to the user Pa based on the interaction record of the user Pa and the engine Aa, a∈{1, 2, 3, …, N} and i≠a.
[0020] In a third aspect of the invention, an AI-based cross-modal personalized recommendation method for comic content is also proposed. This method is applied to a second user, who belongs to a user group comprising multiple users. The method includes the following steps:
[0021] The system receives a first control message and first comic content from a first user in the user group. The first control message contains first modal content. The first comic content is generated by a first AI engine based on the first modal content.
[0022] Based on the first modal content and the first comic content, a second control message is generated, wherein the second control message contains the second modal content;
[0023] The second AI engine recommends second comic content based on the second modality content;
[0024] The first AI engine merges the content of the first comic and the content of the second comic to obtain the target comic.
[0025] The first AI engine is used by the first user, and the second AI engine is used by the second user.
[0026] In a fourth aspect of the present invention, an AI-based cross-modal comic content generation system is also proposed, the system comprising multiple user devices, each user device corresponding to an AI engine;
[0027] The multiple user devices form a user interaction group;
[0028] The first user equipment among the plurality of user equipment is the master equipment, and the other equipment are slave equipment.
[0029] The master device generates first control information, which instructs a first AI engine to generate first comic content based on the first modal content loaded by the first control information, and instructs at least one slave device to respond to the first control information.
[0030] The first AI engine obtains response information from at least one slave device to the first control information, and adjusts the first comic content based on the response information to generate second comic content;
[0031] The response information includes second modality content;
[0032] The first modal content and the second modal content include any one of the following modalities or a combination thereof: text, speech, and image;
[0033] The at least one slave device is specified by the master device based on the semantic analysis results of the first control information.
[0034] In a fifth aspect of the invention, an AI-based cross-modal personalized comic content recommendation system is also proposed. The system includes multiple user devices, each corresponding to an AI engine.
[0035] The multiple user devices form a user interaction group;
[0036] The first user equipment among the plurality of user equipment is the master equipment, and the other equipment are slave equipment.
[0037] The device receives a first control message and first comic content from the main device. The first control message includes first modal content. The first comic content is generated by the first AI engine corresponding to the main device based on the first modal content.
[0038] Based on the first modal content and the first comic content, a second control message is generated from the device, the second control message containing the second modal content;
[0039] The second AI engine corresponding to the device recommends second comic content based on the second modality content.
[0040] The first AI engine merges the content of the first comic and the content of the second comic to obtain the target comic.
[0041] This invention enables creators with different strengths to collaborate efficiently through an innovative collaborative mechanism. A first user generates first control information, loads first modal content, and a first AI engine quickly generates first comic content. Simultaneously, at least one second user can respond to the first control information; their feedback is captured by the first AI engine, which then adjusts the first comic content to generate second comic content. This mechanism makes the entire creative process seamless, eliminating long waiting times for creators at different stages. Creators can work on the work of others at any time, significantly shortening the creation cycle and improving the output efficiency of comic projects. Furthermore, this cross-modal fusion innovation brings unprecedented richness and creativity to comic creation, making the content more engaging and expressive. Finally, based on the comic content generated through group collaboration, the system can also achieve precise personalized recommendations. Because the AI engine collects a large amount of user interaction information and comic content characteristics at different stages during the creation process, through in-depth analysis of this data, the system can accurately understand user preferences and improve the user's comic creation and reading experience.
[0042] Further specific advantages and implementation principles of the present invention will be further detailed in the specific embodiments section in conjunction with the accompanying drawings. Attached Figure Description
[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0044] Figure 1 This is a schematic diagram of the main execution flow of an AI-based cross-modal comic content generation method according to an embodiment of the present invention;
[0045] Figure 2 This is a schematic diagram of the data control flow of an AI-based cross-modal comic content generation method implemented using computer programs;
[0046] Figure 3 This is a schematic diagram illustrating the steps of an embodiment of the AI-based cross-modal comic content personalized recommendation method of the present invention on the second user side.
[0047] Figure 4 This is a schematic diagram of the hardware unit composition of an AI-based cross-modal comic content generation system according to an embodiment of the present invention;
[0048] Figure 5 Based on Figure 4 The flowchart of the implementation of the AI-based cross-modal comic content generation method on the main device side is as follows:
[0049] Figure 6 Based on Figure 4 The flowchart of the implementation of the AI-based cross-modal comic content personalized recommendation method by the system on the device side is described. Detailed Implementation
[0050] In the specific embodiments of this application, if the embodiments of the relevant technical solutions involve user-related data, then when the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0051] In practical implementation, the technical solution of this invention typically involves N users (N>2), who form a collaborative team to complete a comic project (comic generation). Within this team, there is usually a lead person (e.g., a team leader), and the others are collaborative members. Each person (member, user) participates in the entire collaborative process through their own user device.
[0052] Therefore, in various embodiments of the present invention, the terms "user," "member," "user device," "member device," "user terminal," and "member terminal" have the same meaning, and different terms may be used in different contexts. To distinguish between the principal and ordinary members, in subsequent embodiments, the principal (e.g., the group leader) is usually described as "first user," "first user device," "first user terminal," "master device," etc., while other ordinary members are described as "second user," "second user device," "second user terminal," "slave device," etc.
[0053] Meanwhile, related methods or system embodiments may also include descriptions or limitations such as "applied to the first user / second user" or "applied to the master user / master device / slave device." It should be understood that when related methods or system embodiments are described as "applied to the first (user) device" or "applied to the second (user) device," it only means that the related embodiment is described primarily from the perspective of the "first (user) device" or "second (user) device," and does not mean that all steps of that embodiment are performed solely by the first or second device. In fact, unless specifically limited, the implementation process of most method embodiments involves the interaction between the first (user) device and the second (user) device.
[0054] Furthermore, it should be understood that within a team (group), there is only one "first user," "first user device," "first user terminal," or "master device," while there are usually multiple (more than two) "second users," "second user devices," "second user terminals," or "slave devices." The "first user," "first user device," "first user terminal," and "master device" need to interact and collaborate with multiple (more than two) "second user devices," "second user terminals," and "slave devices."
[0055] Based on the above description, the various embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0056] See Figure 1 , Figure 1 This is a schematic diagram of the main execution flow of an AI-based cross-modal comic content generation method according to an embodiment of the present invention.
[0057] exist Figure 1 In this embodiment, the method is applied to a multi-user collaborative group, which includes at least one first user and at least one second user;
[0058] The method includes:
[0059] The first user generates first control information, which instructs the first AI engine to generate first comic content based on the first modal content loaded by the first control information, and instructs at least one second user to respond to the first control information.
[0060] The first AI engine obtains the response information of the second user to the first control information, and adjusts the first comic content based on the response information to generate the second comic content;
[0061] The response information includes second modal content.
[0062] The first modal content and the second modal content include any one of the following modalities or a combination thereof: text, speech, and image.
[0063] For ease of description and in connection with subsequent embodiments, Figure 1 The image shows a user group that collaborates with multiple users. Figure 1 The image shows a group of 5 people, but the actual number could be more or less; the image is for illustrative purposes only. Figure 1 Specifically, users Pi, Pj, and Pk were labeled.
[0064] In the subsequent examples, user Pi is described as the first user (first device / master device), and users Pj and Pk are described as multiple other second users (second devices / slave devices).
[0065] Based on this, see further Figure 2 , Figure 2 This is a schematic diagram of the data control flow of an AI-based cross-modal comic content generation method implemented using computer programs.
[0066] The method is applied to an interactive group containing N users {P1, P2, PN}, where N > 2;
[0067] For ease of description, Figure 2 The flowchart, presented in computer pseudocode, shows the following steps:
[0068] Step 1: User Pi generates control message Ci;
[0069] The control message Ci contains the content of the i-th modality; the content of the i-th modality is taken from any one of the following modalities or a combination thereof: text, voice, and image.
[0070] As an example, the content of the i-th modality can be a text description, such as "In the afterglow of the setting sun, the protagonist stands in front of an ancient castle";
[0071] The control message Ci may also include an instruction message, which specifies that at least one second user (group) should respond to the control message Ci.
[0072] The instruction message can be explicitly specified or implicitly specified;
[0073] Explicit designation can be achieved by using special symbols to remind a specific second user, such as "#userPk" or "@userPj".
[0074] The method of specifying the information can be through semantic attribute limitations, such as "Please have a member skilled in sketching provide further details and submit a first draft of the sketch." In this case, based on the semantic attribute limitation of "members skilled in sketching," the system can automatically filter out "members skilled in sketching" from the team members and remind the relevant member users to respond to the message, assuming that the relevant member users are Pj and Pk;
[0075] Based on the above description, an example of a control message Ci could be...
[0076] "I'm starting a comic project now, with the initial theme being the protagonist standing in front of an ancient castle under the afterglow of the setting sun."
[0077] "Please have members who are skilled in sketching add further details and submit a first draft of the sketch."
[0078] Step 2: Control message Ci will be sent to AI engine Ai, and engine Ai will generate comic content Mi based on the modal content Ti loaded by control message Ci;
[0079] Step 3: Determine the interactive user group InterG based on the control message Ci or the modal content Ti; the interactive user group InterG includes at least user Pj and user Pk; i, j, k ∈ {1, 2, 3, ..., N} and i ≠ j ≠ k;
[0080] In the above example, the control message Ci itself can explicitly or implicitly specify the interactive user group.
[0081] In another embodiment, the control message Ci may not have any semantic attributes that can explicitly or implicitly indicate a specific user group. In this case, the user group InterG can be determined based on the modal content Ti.
[0082] Specifically, when the modal content Ci is the i-th modality, the user group that can respond to other different modalities is identified as the interactive user group;
[0083] Continuing with the example above, if the content of the i-th modality is a text description, then we can identify other users who can generate voice descriptions, image descriptions, or other forms different from text descriptions as the interactive user group. For example, if user Pj is good at voice description input, user Pk is good at sketching first draft input, and users Px and Py are both good at text description input, then we can identify users Pj and Pk as the interactive user group.
[0084] Step 4: The interactive user group (e.g., user Pj and user Pk) generates response information Cj and Ck to control message Ci respectively;
[0085] In one scenario, the response message is input autonomously by users Pj and Pk; in another scenario, the response message is obtained by users Pj and Pk inputting their autonomous modal content into their respective AI engines; in other cases, it can be obtained based on a combination of the two, and this invention does not impose specific limitations on this.
[0086] Preferably, each of the N users {P1, P2, PN} corresponds to an AI engine;
[0087] The response message of user Pa to control message Ci is generated by the AI engine Aa corresponding to user Pa based on the interaction record between user Pa and engine Aa, where a∈{1,2,3,…,N} and i≠a.
[0088] Understandably, even with the same AI engine and the same input, the output will differ depending on the user. This is because the AI engine optimizes and adjusts its output based on the context of the conversation with the current user, resulting in content that better reflects the user's current attributes.
[0089] In implementing embodiments of the present invention, the inventors have noted that the AI engine Aa generates a response message from user Pa to control message Ci based on the interaction record between user Pa and engine Aa.
[0090] In this way, during team collaboration, even if different user members configure the same AI engine, or different members receive the same control message Ci, the corresponding user members will still obtain different modal content when generating response messages based on the control message Ci, thus enriching the creative content of the comic.
[0091] Preferably, different user members can be configured with different AI engines; or, at least some members may not be configured with an AI engine. The AI engine can be any generative AI engine that is already available in the technology, including text-to-text AI engines, text-to-image AI engines, image-to-image AI engines, music generation engines, speech generation engines, etc. The type of AI engine configured by different user members is determined by the content modalities that the user members (devices) support (are good at).
[0092] Optionally, after determining the interactive user group InterG based on the control message Ci or the modal content Ti, the comic content Mi generated by the engine Ai based on the modal content Ti loaded by the control message Ci is also sent to each user member of the interactive user group InterG.
[0093] At this time, each user member of the InterG interactive user group can generate the response message based on the comic content Mi and the control message Ci;
[0094] Step 5: Engine Ai receives and acquires the response information of users Pj and Pk to control message Ci, and adjusts the comic content Mi based on the response information to generate comic content M0.
[0095] As can be seen, in one scenario of the present invention, the at least one second user is designated by the first user based on the first control information, that is, the explicit designation or implicit designation mentioned above;
[0096] In another scenario of the present invention, the at least one second user is automatically determined by the first AI engine based on the first modal content. That is, the control message Ci may not have any semantic attributes that can explicitly or implicitly indicate the designated interactive user group. In this case, the interactive user group InterG can be determined based on the modal content Ti.
[0097] As can be seen, the first specified method has a higher priority than the second automatically determined method.
[0098] As can be seen from the detailed description of the above method embodiments, this invention enables creators with different strengths to work together efficiently through an innovative collaborative mechanism. During multi-person collaboration, creators with different professional expertise oversee the comic content from their respective areas of expertise. Those skilled in textual description ensure the plot is logically coherent and the dialogue is vivid; those skilled in initial comic drafts optimize the composition and character design; and those skilled in music provide musical ideas for creating the comic's atmosphere, ensuring the comic's quality from multiple dimensions. On the other hand, when generating and adjusting comic content, the AI engine can learn from a large amount of comic data, avoiding common creative errors such as inconsistent character designs and chaotic layouts, ultimately creating high-quality, professional-level comic works.
[0099] exist Figures 1-2 Based on the implementation examples of cross-modal comic content generation methods, Figure 3 This is a schematic diagram illustrating the steps of an embodiment of the AI-based cross-modal comic content personalized recommendation method of the present invention on the second user side.
[0100] It should be understood that since the primary user (main device) plays a leading role in team collaboration, they need to exert full initiative and determine the theme, starting point, etc. Therefore, it is not suitable to make recommendations based on AI, otherwise originality cannot be guaranteed.
[0101] However, given that the first user has already determined the theme, location, and other basic elements, other team members can use AI to provide personalized recommendations and guidance to improve work efficiency. Therefore, Figure 3 The personalized recommendation method is mainly used to assist the second user.
[0102] Specifically, Figure 3 The method is applied to a second user who belongs to a user group comprising multiple users, and the method includes the following steps:
[0103] The system receives a first control message and first comic content from a first user in the user group. The first control message contains first modal content. The first comic content is generated by a first AI engine based on the first modal content.
[0104] Based on the first modal content and the first comic content, a second control message is generated, wherein the second control message contains the second modal content;
[0105] The second AI engine recommends second comic content based on the second modality content;
[0106] The first AI engine merges the content of the first comic and the content of the second comic to obtain the target comic.
[0107] Preferably, to avoid the illusion of a large model, Figure 3 The method is further optimized as follows:
[0108] Only receive a first control message from a first user of the user group, the first control message containing first modal content;
[0109] A second control message is generated based on the content of the first modality, and the second control message contains the content of the second modality;
[0110] The second AI engine recommends second comic content based on the second modality content;
[0111] The first AI engine merges the content of the first comic and the content of the second comic to obtain the target comic.
[0112] The first comic content was generated by the first AI engine based on the first modal content.
[0113] In this preferred example, the second user does not accept the original first comic content to avoid duplicate interference.
[0114] It should be understood that from the perspective of the "second user," the "message" generated by the "second user" is its own "second control message"; however, from the perspective of the first user, the "message" generated by the "second user" is a response message to the "control message of the first user."
[0115] In any case, the modality of the "message" (control message) or response message generated by the "second user" is different from that of the control message of the first user. For example, the control message of the first user is a text description message (first modality content), while the "message" (control message) or response message generated by the "second user" is an image description message, such as a sketch or comic draft (second modality content).
[0116] Preferably, different user members can be configured with different AI engines; or, at least some members may not be configured with an AI engine. The AI engine can be any generative AI engine that is already available in the technology, including text-to-text AI engines, text-to-image AI engines, image-to-image AI engines, music generation engines, speech generation engines, etc. The type of AI engine configured by different user members is determined by the content modalities that the user members (devices) support (are good at).
[0117] exist Figures 1-3 Based on the method implementation examples, Figure 4 This diagram illustrates the hardware unit composition of an AI-based cross-modal comic content generation system according to an embodiment of the present invention.
[0118] Figure 4 The illustrated AI-based cross-modal comic content generation system includes multiple user devices, each user device corresponding to an AI engine, and the multiple user devices form a user interaction group; the first user device among the multiple user devices is the master device, and the other devices are slave devices.
[0119] based on Figure 4 , Figure 5 Based on Figure 4 The flowchart of the implementation of the AI-based cross-modal comic content generation method on the main device side is described.
[0120] Specifically, the master device generates first control information, which instructs the first AI engine to generate first comic content based on the first modal content loaded by the first control information, and instructs at least one slave device to respond to the first control information;
[0121] The first AI engine obtains response information from at least one slave device to the first control information, and adjusts the first comic content based on the response information to generate second comic content;
[0122] The response information includes second modality content;
[0123] The first modal content and the second modal content include any one of the following modalities or a combination thereof: text, speech, and image;
[0124] The at least one slave device is determined by the master device based on the semantic analysis results of the first control information.
[0125] Specifically, the at least one slave device is determined by the master device based on the semantic analysis results of the first control information, including one of the following methods:
[0126] In one scenario of the present invention, the at least one second user is designated by the first user based on the first control information, i.e., the explicit or implicit designation mentioned above.
[0127] In another scenario of the present invention, the at least one second user is automatically determined by the first AI engine based on the first modal content. That is, the control message Ci may not have any semantic attributes that can explicitly or implicitly indicate the designated interactive user group. In this case, the interactive user group InterG can be determined based on the modal content Ti.
[0128] Further embodiments of the present invention can also be a cross-modal personalized comic content recommendation system based on AI, the system including multiple user devices, each user device corresponding to an AI engine, and the system structure is similar to... Figure 4 similar.
[0129] Figure 6 Based on Figure 4 The implementation flowchart of the AI-based cross-modal comic content personalized recommendation method implemented by the system on the device side specifically includes:
[0130] The device receives a first control message and first comic content from the main device. The first control message includes first modal content. The first comic content is generated by the first AI engine corresponding to the main device based on the first modal content.
[0131] Based on the first modal content and the first comic content, a second control message is generated from the device, the second control message containing the second modal content;
[0132] The second AI engine corresponding to the device recommends second comic content based on the second modality content.
[0133] The first AI engine merges the content of the first comic and the content of the second comic to obtain the target comic.
[0134] Preferably, to avoid the illusion of a large model, Figure 6The method is further optimized as follows:
[0135] The slave device receives only a first control message from the master device, the first control message containing first modal content;
[0136] Based on the first modal content and the first comic content, a second control message is generated from the device, the second control message containing the second modal content;
[0137] The second AI engine corresponding to the device recommends second comic content based on the second modality content.
[0138] The first AI engine merges the first comic content and the second comic content to obtain the target comic; the first comic content is generated by the first AI engine corresponding to the main device based on the first modal content.
[0139] In this preferred example, the device does not accept the original first comic content to avoid duplicate interference.
[0140] It should be understood that from the perspective of the "slave device", the "message" generated by the "slave device" is its own "secondary control message"; however, from the perspective of the master device, the "message" generated by the "slave device" is a response message to the "master device's control message".
[0141] Regardless of the specific modality, the "messages" (control messages) or response messages generated by the "slave device" are different from the control messages of the master device. For example, the control messages of the master device are text description messages (first modality content), while the "messages" (control messages) or response messages generated by the "slave device" are image description messages, such as sketches or comic drafts (second modality content).
[0142] Preferably, different slave devices can be configured with different AI engines; or, at least some slave devices are not configured with AI engines. The AI engine can be various generative AI engines that are already available in the prior art, including text-to-text AI engines, text-to-image AI engines, image-to-image AI engines, music generation engines, speech generation engines, etc. The type of AI engine configured for different slave devices (users / members) is determined by the content modalities that the slave device (user) supports (is good at).
[0143] During team collaboration, the AI engines on these different devices do not work in isolation but rather in close coordination. Taking a comic book project creation process as an example, the creator responsible for the text description first uses a text processing AI engine on their device to generate detailed story scripts and character settings—the initial control information. This information is then shared with the team collaboration platform. Next, the creator skilled in initial comic drafts retrieves information from the platform and uses their device's image generation AI engine to generate the first version of the comic's visuals. At this point, other team members, such as creators skilled in music, can use the generated comic visuals and text script on their own devices, leveraging an audio AI engine, to begin conceiving and generating preliminary music and sound effects schemes. Simultaneously, team members can also provide multimodal interactive feedback on the generated comic content using their respective devices' AI engines. For example, through voice interaction, they can directly suggest modifications to the character expressions and action details in the visuals; or using handwriting input, they can directly mark the areas and content requiring adjustment on the image. The AI engine will then quickly optimize and adjust the comic content based on this feedback, generating a more complete second version of the comic.
[0144] Although not shown in the accompanying drawings, a preferred and more common product embodiment may also be an electronic device comprising: a memory and one or more processors. The memory stores one or more applications adapted to be executed by the one or more processors, representing the aforementioned AI-based cross-modal comic content generation method or personalized recommendation method.
[0145] Although not shown in the accompanying drawings, further embodiments also include a computer-readable storage medium storing a computer program that, when executed, implements the steps of the aforementioned personalized recommendation method.
[0146] It is understood that the system, product, equipment, and media implementation examples and method implementations correspond to each other and can be referenced by each other, and their principles are similar or the same, so they will not be elaborated again.
[0147] Other technologies, principles, algorithms, or models not elaborated in detail in this application can be found in the prior art.
[0148] The technical solution proposed in this invention is based on a team collaboration model where different devices utilize different AI engines to output multimodal comic content, offering significant advantages over traditional comic creation. Firstly, it breaks down the limitations of time and space; team members, regardless of their location, can engage in creative work using their respective AI engines as long as they access the collaboration platform through their devices, greatly improving the flexibility of collaboration. Secondly, different AI engines leverage their expertise in their respective areas, achieving deep integration and efficient generation of multimodal content, significantly improving the quality and efficiency of comic creation, and bringing unprecedented innovative development opportunities to the comic industry.
[0149] In this invention, multiple user devices correspond to different AI engines, each capable of deep understanding and fusion of multimodal content such as text, images, and music. When generating comic content, textual descriptions can be accurately transformed into vivid images, and elements within the images can be flexibly adjusted according to the text's requirements. This cross-modal fusion innovation brings unprecedented richness and creativity to comic creation, making comic content more engaging and expressive.
[0150] Based on comic content generated through group collaboration, this invention also enables precise personalized recommendations. During the creation process, the AI engine collects a large amount of user interaction information and comic content characteristics at different stages. Through in-depth analysis of this data, the system can accurately understand user preferences. For example, if a user group repeatedly selects a specific style of character design and plot direction during creation, the system can accurately identify this group's preferences and recommend comics with similar styles and plot structures. Alternatively, in subsequent creations, the system can provide creators with more tailored creative suggestions based on previous preferences, such as recommending relevant image materials and music styles, thereby enhancing the user's comic creation and reading experience.
[0151] From the perspective of comic quality, this invention has significant advantages. On the one hand, in the collaborative process, creators with different professional expertise oversee the comic content from their respective areas of expertise. Those skilled in textual description ensure the plot is logically coherent and the dialogue is vivid; those skilled in initial drafts optimize the composition and character design; and those skilled in music provide musical ideas for creating the comic's atmosphere, ensuring the comic's quality from multiple dimensions. On the other hand, when generating and adjusting comic content, the AI engine can learn from a large amount of comic data, avoiding common creative errors such as inconsistent character designs and chaotic layouts, ultimately creating high-quality, professional-level comic works.
[0152] The foregoing has shown and described the method embodiments and systems of the present invention, but it will be understood by those skilled in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An AI-based cross-modal cartoon content generation method, characterized in that, the method is applied to a user group of multiple people cooperating, which includes at least one first user and at least one second user; the method comprises: generating first control information by the first user, the first control information indicating that a first AI engine generates first cartoon content based on first modal content loaded by the first control information, and at least one second user responding to the first control information; the first AI engine acquires response information of the second user to the first control information, and adjusts the first cartoon content based on the response information to generate second cartoon content; wherein the response information includes second modal content.
2. The AI-based cross-modal cartoon content generation method of claim 1, characterized in that, the first modal content and the second modal content include any one of the following modalities or a combination thereof: text, voice, image.
3. The AI-based cross-modal cartoon content generation method of claim 1, characterized in that, the at least one second user is specified by the first user based on the first control information.
4. The AI-based cross-modal cartoon content generation method of claim 1, characterized in that, the at least one second user is automatically determined by the first AI engine based on the first modal content.
5. An AI-based cross-modal cartoon content generation method, characterized in that, The method is applied to an interactive group containing N users {P1, P2, …, PN}, N>2, and the method comprises the following steps: When user Pi generates control message Ci, send the control message Ci to AI engine Ai, and engine Ai generates cartoon content Mi based on modal content Ti loaded by control message Ci; determine the interactive user group InterG based on the control message Ci or the modal content Ti; the interactive user group InterG at least contains user Pj and user Pk; i, j, k ∈ {1, 2, 3, …, N} and i≠j≠k; engine Ai receives and acquires response information of user Pj and user Pk to control message Ci, and adjusts the cartoon content Mi based on the response information to generate cartoon content M0; each of the N users {P1, P2, …, PN} corresponds to an AI engine; The response message of user Pa to control message Ci is generated by AI engine Aa corresponding to user Pa based on the interaction record of user Pa and engine Aa, a ∈ {1, 2, 3, …, N} and i≠a.
6. An AI-based cross-modal cartoon content personalized recommendation method, the method being applied to a second user, the second user belonging to a user group comprising a plurality of users, characterized in that, The method comprises the following steps: receive the first control message and the first cartoon content from the first user of the user group, the first control message containing first modal content; the first cartoon content is generated by the first AI engine based on the first modal content; generate a second control message based on the first modal content and the first cartoon content, the second control message containing second modal content; recommend second cartoon content by the second AI engine based on the second modal content; The first AI engine fuses the first comic content and the second comic content to obtain a target comic.
7. The AI-based cross-modal comic content personalized recommendation method of claim 6, wherein, The first AI engine is used by the first user, and the second AI engine is used by the second user.
8. An AI-based cross-modal comic content generation system, the system comprising a plurality of user devices, each user device corresponding to an AI engine, characterized in that: The plurality of user devices form a user interaction group. A first user device in the plurality of user devices is a master device, and other devices are slave devices. The first control information is generated by the master device, and indicates that the first AI engine generates first comic content based on first modal content loaded by the first control information, and at least one slave device responds to the first control information; The first AI engine obtains response information of at least one slave device to the first control information, and adjusts the first comic content based on the response information to generate second comic content; The response information includes second modal content; The first modal content and the second modal content include any one of the following modalities or a combination thereof: text, voice, image; The at least one slave device is specified by the master device based on a semantic analysis result of the first control information.
9. An AI-based cross-modal comic content personalized recommendation system, the system comprising a plurality of user devices, each user device corresponding to an AI engine, characterized in that: The plurality of user devices form a user interaction group. A first user device in the plurality of user devices is a master device, and other devices are slave devices. The slave device receives a first control message and first comic content from the master device, the first control message containing first modal content; the first comic content is generated by the first AI engine corresponding to the master device based on the first modal content; Based on the first modal content and the first comic content, the slave device generates a second control message containing second modal content; The second AI engine corresponding to the slave device recommends second comic content based on the second modal content; The first AI engine fuses the first comic content and the second comic content to obtain a target comic.
Citation Information
Patent Citations
Material-engine-based automatic cartoon generating method
CN102110304B
Cartoon image generation method and device, equipment and storage medium
CN119810223A
Interactive cartoon generation system and method based on generative AI technology and storage medium
CN117391923A
Image-text generation method and system based on AR shooting and artificial intelligence
CN118247392A