Media information generation method and apparatus, electronic device, storage medium, and product

By extracting multi-dimensional structured features and adjusting intelligent systems, media information is generated, solving the problems of limited creative forms and insufficient creative diversity in existing technologies, and realizing the generation of media information with creative diversity and differentiation.

CN122120576APending Publication Date: 2026-05-29DOUYIN VISION CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DOUYIN VISION CO LTD
Filing Date
2026-03-31
Publication Date
2026-05-29

Smart Images

  • Figure CN122120576A_ABST
    Figure CN122120576A_ABST
Patent Text Reader

Abstract

The present application provides a media information generation method, device, electronic equipment, storage medium and product. The method comprises: in response to a media information generation task request, determining first media information; generating second media information based on the first media information, the second media information being media information obtained by at least one intelligent system based on content generation of a feature description content corresponding to the first information, the first information being obtained by adjusting a feature description content corresponding to the second information by the at least one intelligent system, and the at least one intelligent system adjusting the feature description content corresponding to the second information according to a first target, the second information being determined based on a plurality of dimensions for extracting structured feature description content of the first media information, and the first target indicating that a feedback effect of media information generated based on the first information is not lower than a feedback effect of media information generated based on the second information. The present application solves the problem of single creative form and poor creative diversity in media information generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This article relates to the field of computer processing technology, and in particular to a method, apparatus, electronic device, storage medium and product for generating media information. Background Technology

[0002] With the rapid development of content generation technology, its application in the field of media information generation is becoming increasingly widespread, and various multimedia platforms can use related technologies to improve the efficiency of media information generation.

[0003] In related solutions, media information can be generated in batches by filling in information elements based on preset templates and fixed rules. While this method has high generation efficiency, the generated media information suffers from limited creative forms and poor creative diversity. Alternatively, media information can be generated in batches using generative models, which improves creative diversity to some extent, but the generation process relies on human experience, resulting in limited creative diversity in the final media information and failing to meet practical application needs. Summary of the Invention

[0004] This paper provides a method, device, electronic device, storage medium, and product for generating media information, which solves the problems of limited creative forms and poor creative diversity in media information generation.

[0005] In one scenario, this paper provides a method for generating media information, the method comprising: In response to the media information generation task request, determine the first media information; Second media information is generated based on the first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined based on structured feature description content extracted from the first media information from multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0006] In one instance, this document also provides a media information generation apparatus, the apparatus comprising: The acquisition module is used to determine the first media information in response to the media information generation task request; A generation module is used to generate second media information based on the first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined based on structured feature description content extracted from the first media information from multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0007] In one instance, this document also provides an electronic device comprising: One or more processors; Storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the media information generation method as described herein.

[0008] In one instance, this document also provides a storage medium containing computer-executable instructions that, when executed by a computer processor, are used to perform the media information generation method as described herein.

[0009] In another scenario, this document also provides a computer program product, including a computer program that, when executed by a processor, implements the media information generation method as described herein.

[0010] In the above method, structured feature extraction of the first media information from multiple dimensions yields the second information. This method clearly identifies the feature descriptions of different dimensions within the first media information. Then, at least one intelligent system adjusts and optimizes the feature descriptions corresponding to the second information to obtain the first information, with the first objective indicating that the feedback effect of the media information generated based on the first information is no less than that generated based on the second information. Since the feature descriptions corresponding to the second information are multi-dimensional, discrete, and independently adjustable, at least one intelligent system can flexibly adjust and optimize them while maintaining the feedback effect, allowing the generated feature descriptions corresponding to the first information to produce better feedback. This effectively alleviates the problems of homogenization and insufficient creative diversity in the media information generation process. Because the feature descriptions corresponding to the first information are structured feature descriptions optimized based on feedback effects, the second media information generated by at least one intelligent system based on these feature descriptions exhibits differences in presentation, visual style, and copywriting. This addresses the problems of singular creative forms and insufficient creative diversity in the media information generation process while ensuring the feedback effect of the generated second media information is not reduced as much as possible. Attached Figure Description

[0011] The above and other features, advantages, and aspects of the embodiments described herein will become more apparent when taken in conjunction with the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale.

[0012] Figure 1 This is a schematic diagram of the structure of a media information generation system under one scenario. Figure 2 This is a flowchart illustrating a media information generation method in one scenario. Figure 3 This is a flowchart illustrating a media information generation method in another scenario. Figure 4 A schematic diagram of the system architecture of at least one intelligent system under one scenario; Figure 5 This is a schematic diagram illustrating the optimization and reinforcement learning of multi-dimensional structured features in media information under a specific scenario. Figure 6 This is a schematic diagram of a media information generation device in one scenario. Figure 7 This is a schematic diagram of the structure of an electronic device that implements a media information generation method in one scenario. Detailed Implementation

[0013] The embodiments will now be described in more detail with reference to the accompanying drawings. While some embodiments are shown in the drawings, it should be understood that the technical solution can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present technical solution. It should be understood that the illustrated drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the technical solution.

[0014] It should be understood that the steps described in the method implementation may be performed in different orders and / or in parallel. Furthermore, the method implementation may include additional steps and / or omit the steps shown. The scope of this document is not limited in this respect.

[0015] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one situation" means "at least one situation"; the term "another situation" means "at least one additional situation"; the term "some situations" means "at least some situations". Definitions of other terms will be given in the following description.

[0016] It should be noted that the concepts of "first" and "second" mentioned in this article are only used to distinguish different devices, modules or units, and are not used to limit the order of the functions performed by these devices, modules or units or their interdependencies.

[0017] It should be noted that the terms "one" and "more" used in this document are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0018] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0019] It is understood that before using the technical solutions disclosed in the various embodiments of this document, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this document in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0020] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as electronic devices, applications, servers, or storage media, that perform the operations described herein, based on the prompt message.

[0021] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0022] It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation method described in this article. Other methods that comply with relevant laws and regulations may also be applied to the implementation method described in this article.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] The technical solution in this article can be applied to... Figure 1 The media information generation system shown is illustrated. In practical applications, this media information generation system may include a client 101 and a server 102. The client 101 may include, but is not limited to, browsers, applications (Apps), Hyper Text Markup Language (HTML) applications, lightweight applications (also known as mini-programs), or cloud applications. The client 101 may be deployed on an electronic device, relying on the operation of that device or certain applications within the device to implement its functions. The electronic device may be, for example, a device with a display screen that supports information browsing, such as a smartphone, tablet, personal computer, or other client terminal. For ease of understanding, Figure 1 The client is primarily represented in the form of a device. Other applications, such as media information generation applications, can also be configured on the electronic device. Server 102 can be one or more servers providing various services. That is, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server; furthermore, it can be a server for a distributed system, a server integrating blockchain technology, a cloud server, or an intelligent cloud computing server or intelligent cloud host deployed with machine learning models, etc.

[0025] It should be noted that in this technical solution, the media information generation method can be executed by client 101, or by client 101 and server 102, with different functional parts of the corresponding media information generation device deployed on client 101 and server 102 respectively; wherein, the first processing module, second processing module, and third processing module of the device are deployed on client 101. It should be understood that... Figure 1The number of clients and servers shown is for illustrative purposes only. Any number of clients and servers can be configured to meet specific implementation requirements.

[0026] Figure 2 A flowchart illustrating a media information generation method is provided. This method is applicable to media information generation scenarios requiring both large-scale and differentiated generation, where the generated media information must balance feedback effectiveness and creative diversity. The method can be executed by a media information generation device, which can be implemented through software and / or hardware, optionally through electronic devices such as mobile terminals, PCs, or servers. Figure 2 As shown, the media information generation method in this paper may specifically include the following processes: S210. In response to the media information generation task request, determine the first media information.

[0027] A media information generation task request can be an instruction request used to trigger the start of the media information generation process. The media information generation task request carries the corresponding task requirements, task objectives, and pre-configured constraints for media information generation. Media information generation task requests can be generated through external active triggering or through system-timed triggering; both methods can initiate media information generation.

[0028] Media information generation task requests include, but are not limited to, the following: Media information generation task requests can be instructions for generating text and image-based media information, used to generate media information adapted to specified scenarios for promotion and display; Media information generation task requests can be requests for optimizing media information, used to initiate optimization and update tasks for already generated media information, generating media information with differentiated characteristics; Media information generation task requests can be instructions for batch generating media information, used to specify the quantity, content dimensions, and quality standards of the media information to be generated.

[0029] Media information can be digital information content presented on digital terminals for information display and dissemination. Media information can be used, displayed, or disseminated independently. Media information can include at least one of the following forms of expression: text, images, text-image combinations, and video. Media information can include, but is not limited to, the following types: image-based media information, text-based media information, text-image combination media information, video-based media information, and multimodal composite media information. Among them, image-based media information includes advertising images, promotional posters, product display images, information flow advertising images, cover images, and graphic creative materials; text-based media information includes advertising headlines, promotional copy, advertising slogans, product descriptions, event copy, and guiding copy; text-image combination media information includes complete advertising materials, promotional materials, promotional flyers, etc., formed by combining text and images; video-based media information includes short advertising videos, promotional videos, dynamic creative videos, and interstitial advertising videos.

[0030] In some cases, it can be combined with one or more of the above-mentioned alternatives. The first media information can be media information matched and filtered through an external or built-in media information recommendation system. The first media information can also be created by the user. Specifically, the first media information includes media information obtained through at least one of the following methods: drawing, editing, custom editing, and local generation. For example, self-drawn pictures, self-written copy, self-designed posters, and custom-made creative materials.

[0031] In some cases, the above-mentioned alternative solutions can be combined to determine the first media information in response to a media information generation task request. This includes: displaying multiple media information items on a media information page in response to the media information generation task request; and determining the at least one media information item selected by the selection operation as the first media information in response to a selection operation on at least one media information item among the multiple media information items. The multiple media information items displayed on the media information page are media information matched and filtered through a media information recommendation system.

[0032] In some cases, the above-mentioned alternative solutions can be combined to determine the first media information in response to a media information generation task request. This includes: displaying multiple media information items on a media information page in response to the media information generation task request; presenting an editing page for at least one media information item in response to an edit trigger operation for at least one of the multiple media information items; generating an editing result for at least one media information item in response to performing an editing operation on at least one media information item; and using the editing result of at least one media information item as the first media information item. The multiple media information items displayed on the media information page are media information matched and filtered through a media information recommendation system.

[0033] In some cases, the above-mentioned alternative solutions can be combined to determine the first media information in response to a media information generation task request. This includes: displaying multiple media information items on a media information page in response to the media information generation task request; presenting an editing page for at least one media information item in response to an edit trigger operation for at least one of the multiple media information items; generating an editing result for at least one media information item in response to performing an editing operation on at least one media information item; and using the editing result of at least one media information item as the first media information. The multiple media information items displayed on the media information page can be media information obtained through drawing, secondary editing, or custom editing.

[0034] The technical effect of the above method is that by triggering the media information generation task request, the media information generation process can be started on demand, adapting to the need for real-time media information generation in various scenarios; the baseline first media information is obtained simultaneously, providing a reference for subsequent structured feature decomposition and effect-oriented optimization, avoiding the lack of reference basis for the optimization process.

[0035] S220. Generate second media information based on first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined by extracting structured feature description content from the first media information based on multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0036] The second information can be multiple feature descriptions obtained after extracting structured features from the first media information from multiple dimensions. These feature descriptions can be represented using keywords or vectors. The feature descriptions corresponding to the second information can correspond to the following dimensions: visual style, composition, elements, text content, color scheme, and layout. The feature descriptions corresponding to the second information possess the characteristics of being discretizable, independently adjustable, and interpretable.

[0037] This paper pre-configures at least one intelligent system, which refers to a system capable of autonomous control based on a machine learning model. An intelligent system is, for example, a virtual object or entity capable of making decisions and autonomously executing actions based on a machine learning model to achieve preset goals or complete preset tasks. An intelligent system can be an automated program that understands the user's intent and can utilize models or invoke tools to complete various types of tasks. In some cases, examples of intelligent systems may include, but are not limited to: agents, bots, chatbots, digital avatars, intelligent customer service, digital assistants, etc. Alternatively, an intelligent system can also be an intelligent role implemented based on a machine learning model. An "intelligent system" can process user requests based on generative models (e.g., language models, multimodal models) to perform specified types of tasks. In some cases, an intelligent system may also relate to virtual accounts, which may have corresponding avatars or nicknames.

[0038] After determining the second information, at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The first objective indicates that the feedback effect of media information generated based on the first information is not lower than the feedback effect of media information generated based on the second information. The feature description content corresponding to the second information is adjusted and optimized according to the first objective to obtain the first information, so that the feedback effect of using media information generated based on the first information is not lower than the feedback effect of using media information generated based on the second information. The feedback effect of using media information can be characterized by at least one of the following indicators: exposure, click-through rate, conversion rate, interaction rate, and completion rate.

[0039] The first information can be structured feature data optimized and adjusted by at least one intelligent system based on the feedback effect after using the media information generated from the first information. It represents the optimized and upgraded feature data of the corresponding feature descriptions of the second information, while retaining feature descriptions that enable better feedback after using the media information. Therefore, the second media information can be generated by at least one intelligent system based on the feedback-oriented optimized feature descriptions of the first information, and can be used for media information use and display.

[0040] The technical effect of the above method is that by extracting multi-dimensional structured features from the first media information, second information that can be parsed, split, and independently adjusted is obtained. Furthermore, with the goal of ensuring that the feedback effect of media information generated based on the first information is no less than that of media information generated based on the second information, at least one intelligent system optimizes and adjusts the second information to obtain the first information. By introducing the feedback effect of the media information, at least one intelligent system is assisted in adjusting and optimizing the corresponding feature descriptions of the second information in a direction that produces better feedback effects. Then, based on the first information, at least one intelligent system generates second media information. This approach can ensure that the feedback effect of the media information is not reduced, while allowing the generated second media information to differ in presentation, visual style, and copywriting, effectively solving the problems of creative homogenization, insufficient creative diversity, and insufficient creative controllability in media information generation.

[0041] In some cases, the first and second information can be represented in natural language, in combination with one or more of the alternative options mentioned above. Further alternatively, the first and second information can be represented in explicit, structured, and interpretable natural language.

[0042] Both the first and second pieces of information are presented in explicit, structured, and interpretable natural language. In this case, they are presented as directly readable text statements, keywords, or logical expressions, possessing clear grammatical structures and semantic relationships. For example, the feature description corresponding to the first piece of information can be directly reflected in specific elements such as "red main color" and "large-font promotional headline"; or, for example, the feature description corresponding to the first piece of information can be directly reflected in {style: retro, scene: coffee shop, elements: [latte, book], color tone: warm, composition: close-up}. It is evident that the presentation of the first and second pieces of information in this scenario is highly interpretable, intuitively reflecting the constituent elements of the media information. The adjustment and optimization process from the feature description corresponding to the second piece of information to that corresponding to the first piece of information is transparent, facilitating manual review, intervention, and strategy adjustments. Furthermore, the presentation of the first piece of information in this scenario provides precise and discrete control over the feature description content, ensuring the controllability and compliance of the media information generation process. Representing the first and second information in explicit, structured, and interpretable natural language not only makes unstructured media information analyzable and attributable, but also provides clear and unambiguous instructions for the media information generation process. This allows the optimized generation of media information to be directly transformed into a quantitative operation of the structured feature description content, and then uses the optimized structured feature description content to generate optimized media information, greatly simplifying the complexity of cross-modal media information optimization.

[0043] The technical effect of the above method is that by representing the first and second information in an explicit, structured, and interpretable natural language form, multimodal media information can be mapped into a standardized text description system. At this time, the first and second information also have the feature understanding output of the media information, providing a structured and interpretable intermediate representation for the media information. This can break down the modal barriers of multimodal content and unify different forms of media information such as images, copywriting, and posters into a standardized description that can be intuitively understood. This simplifies the complex cross-modal optimization problem into a strategy optimization problem of combining feature description content. The optimization process becomes transparent and traceable, which not only facilitates problem attribution but also provides data support for more refined feature description content optimization strategy control.

[0044] In some cases, the first and second information can be represented in vector feature form, combining various alternative schemes from one or more of the above scenarios. Further alternatively, the first and second information can be represented in implicit and non-interpretable vector feature form.

[0045] Both the first and second information are implicit and uninterpretable vector features. In this form, the first and second information are encoded as high-dimensional or low-dimensional real vectors. The numerical values ​​in the vector space are merely intermediate states of the model's calculations and do not possess direct semantic correspondences. They cannot be directly interpreted as specific visual or textual elements. In this case, the presentation of the first and second information utilizes deep neural networks to capture deep relationships between data, possessing powerful feature extraction and generalization generation capabilities. However, the feature adjustment process is presented as a black-box numerical optimization, lacking direct semantic interpretation.

[0046] When the first information and the second information are represented in implicit and uninterpretable vector feature form, the first media information can be encoded into a high-dimensional vector. The feature description content corresponding to the second information represented in vector form can be extracted from the first media information. Then, the first information represented in vector form can be obtained by adjusting the feature description content corresponding to the second information represented in vector form. Finally, the first information represented in vector form is used to implicitly guide the generation of the second media information.

[0047] The technical advantages of the above approach are that by representing the first and second information in an implicit, uninterpretable real-number vector form, efficient encoding can be achieved through deep neural networks, fully capturing the deep data relationships of multimodal media information, and mining potential features that are difficult to obtain through conventional structured analysis, thus possessing stronger feature extraction and generalization generation capabilities. Moreover, by encoding the first media information into a high-dimensional vector, the second information in vector form can be extracted quickly, and the feature adjustment of the vector dimension and the generation of the first information can be completed efficiently. The entire process is guided by implicit numerical calculations of the model, making the generation process simpler and the computational efficiency higher. It is compatible with end-to-end intelligent generation architecture and can quickly output diverse media information, making up for the limitations of explicit labeling in deep feature mining.

[0048] In some cases, it may be combined with various alternatives from one or more of the above situations, with at least one intelligent system being a third model. The third model includes a machine learning model for receiving, processing, and / or generating at least two different modalities of information. Generating second media information based on first media information may include the following steps: Obtain first prompt information, which includes third information, fifth information, first task objective, and sixth information. The first prompt information is associated with the first task objective, which indicates the task objective obtained from parsing the media information generation task request. The third information is obtained by extracting structured feature descriptions of the third media information based on multiple dimensions. The fifth information is used to indicate the feedback effect of the third media information associated with the third information. The sixth information is used to indicate the adjustment and optimization strategy for the feature descriptions corresponding to the second information. The third media information is the media information that has been used and has produced feedback effects before triggering the media information generation task request. Generate second media information based on the first prompt information, the third model, and the first media information.

[0049] The technical solution described above utilizes a powerful and unified multimodal machine learning model to complete the media information generation task of generating second media information based on first media information. Its working method involves constructing a complex first prompt message through prompt engineering. This prompt message includes third information (containing structured feature descriptions extracted from third media information across multiple dimensions), feedback effects of the associated third media information, adjustment and optimization strategies for the feature descriptions corresponding to the second information, and the task objective parsed from the media information generation task request. This complex first prompt message is then directly input into a third model, with the expectation that the third model, guided by the first prompt message, can directly understand all the context and generate the final optimized second media information within one or a few interactions.

[0050] The technical effect of this solution is that it extracts structured features from first media information from multiple dimensions to obtain second information. It can clearly identify the feature descriptions of different dimensions within the first media information. Then, through at least one intelligent system, it can adjust and optimize the feature descriptions corresponding to the second information according to a first objective, obtaining the first information. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than that generated based on the second information. Because the feature descriptions corresponding to the second information are multi-dimensional, discrete, and independently adjustable, at least one intelligent system can flexibly adjust and optimize them while maintaining the feedback effect, so that the generated feature descriptions corresponding to the first information can produce a better feedback effect, effectively alleviating the problems of homogenization and insufficient creative diversity in the media information generation process. Since the feature descriptions corresponding to the first information are structured feature descriptions optimized and adjusted based on feedback effect, the second media information generated by at least one intelligent system based on the feature descriptions corresponding to the first information exhibits differences in presentation, visual style, and copywriting. This solves the problems of singular creative forms and insufficient creative diversity in the media information generation process while ensuring that the feedback effect of the generated second media information is not reduced as much as possible.

[0051] Figure 3 A flowchart illustrating another method for generating media information is provided. The technical solution in this scenario further optimizes the process of generating second media information based on first media information in the aforementioned scenarios. The technical solution in this scenario can be combined with various optional solutions in one or more of the aforementioned scenarios.

[0052] like Figure 3 As shown, the media information generation method provided in this paper may include the following processes: S310. In response to the media information generation task request, determine the first media information.

[0053] S320. Based on the first intelligent system, a first task objective and a first execution order are determined. The first execution order is associated with the first task objective. The first task objective indicates the task objective obtained by parsing the task request generated from the media information. The first execution order is used to call at least two second intelligent systems and instruct the two second intelligent systems to collaboratively execute their respective associated tasks. The first intelligent system is also used to detect the execution status of the associated tasks of the at least two second intelligent systems. At least one intelligent system includes the first intelligent system and at least two second intelligent systems.

[0054] See Figure 4At least one intelligent system includes a first intelligent system. The first intelligent system is an intelligent system used for scheduling in a distributed collaborative architecture. It has the functions of task parsing, target decomposition, process planning, collaborative scheduling and status monitoring. It is different from the intelligent system that executes specific tasks. The first intelligent system is used to coordinate the overall tasks, rather than directly performing feature extraction, content generation and media information generation.

[0055] See Figure 4 At least one intelligent system also includes at least two second intelligent systems. For example, the at least two second intelligent systems may include a third, fourth, and fifth intelligent system. The at least two second intelligent systems are intelligent systems in a distributed collaborative architecture used to execute specific tasks. Each intelligent system in the at least two second intelligent systems has a clear division of labor and a specialized function. They respectively undertake the task of extracting structured feature description content from the first media information, the task of optimizing the feature description content corresponding to the second information, and the task of generating content based on the feature description content corresponding to the first information. They have the ability to independently execute their corresponding tasks, have no global scheduling authority, and only need to complete their own tasks according to the scheduling instructions of the first intelligent system. The boundaries of each intelligent system in the at least two second intelligent systems are clear, which facilitates maintenance, iteration, and functional expansion.

[0056] The primary task objective can be the task requirements extracted from the externally input media information generation task request after parsing. The primary task objective can include the type of media information to be generated, the usage scenario of the media information, the compliance requirements of the media information, the lower limit of the feedback effect of the media information, and the style of the media information; it serves as the task guide for the execution of the media information generation task.

[0057] The first execution order is used to invoke at least two secondary intelligent systems and instruct the two secondary intelligent systems to collaboratively execute their respective associated tasks in sequence. The first execution order can be a task execution sequence adapted to the at least two secondary intelligent systems, planned and formulated by the first intelligent system according to the first task objective. It can be used to indicate the start time, execution logic, and connection nodes of each intelligent system in the at least two secondary intelligent systems, while limiting the collaborative relationship between the intelligent systems in the at least two secondary intelligent systems. This avoids conflicts, duplicate executions, or process breaks in the tasks to be executed by the intelligent systems in the at least two secondary intelligent systems, ensuring the orderly progress and seamless connection of subdivided tasks.

[0058] The execution status of the tasks associated with the second intelligent system can be used to indicate the execution progress, completion status, and abnormal status of the corresponding tasks in the second intelligent system. By observing the execution status of the tasks associated with the second intelligent system, it is possible to promptly identify whether the corresponding tasks in the second intelligent system are experiencing problems such as task lag, execution failure, or unsatisfactory results. This enables closed-loop process control capabilities, ensuring that the media information generation task proceeds stably according to the first execution order and avoiding the interruption of the overall media information generation task process due to the failure of some intelligent agents in at least two second intelligent systems.

[0059] The technical effect of the above approach is that it enables overall task scheduling through the first intelligent system, breaks down complex media information generation tasks into subdivided tasks, and performs unified arrangement and collaboration through the first intelligent system, clarifying the collaboration order of each intelligent system in at least two second intelligent systems; it ensures the stability of the closed-loop process by real-time task execution status of each intelligent system in at least two second intelligent systems, achieves clear task division and clear responsibilities, and greatly improves the controllability and stability of the overall media information generation process.

[0060] S330. According to the first execution order, second media information is generated based on first media information and at least two second intelligent systems. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined based on the structured feature description content extracted from the first media information from multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0061] See Figure 4 In the scheduling and initiation phase of the collaborative process of at least two second intelligent agents, the first intelligent system acts as the scheduling center. It first connects with media information to generate task requests, completes request parsing, and extracts the first task objective. Then, based on the first task objective, it generates a first execution order that matches the first task objective, clarifying the collaborative logic and execution order of each intelligent system in at least two second intelligent systems. At the same time, according to the first execution order, it initiates the execution status detection of the tasks associated with each intelligent system in at least two second intelligent systems, tracks the task execution status of each intelligent system in at least two second intelligent systems throughout the process, realizes unified control and scheduling of global tasks, and ensures orderly connection of each link.

[0062] The technical effect of the above method is that the multi-agent execution of each intelligent system in at least two second intelligent systems is coordinated, the task is subdivided and processed in a specialized manner, and the extraction of feature description content and the accuracy of content generation are improved; the feature description content corresponding to the second information is optimized without reducing the effect of the generated media information, taking into account the feedback effect and the creative diversity of media information, solving the problems of creative homogenization and insufficient creative diversity, and the multi-agent execution of each intelligent system in at least two second intelligent systems can improve the generation efficiency of media information.

[0063] In some cases, it can be combined with one or more of the alternative solutions mentioned above to generate second media information based on the first media information and at least two second intelligent systems, including the following steps 1a-3a: Step 1a: Generate second information based on the third intelligent system and the first media information. The third intelligent system is an intelligent system among at least two second intelligent systems. The third intelligent system is used to transform unstructured media information into structured feature description content.

[0064] Step 2a: Generate first information based on the fourth intelligent system and the second information. The fourth intelligent system is an intelligent system among at least two second intelligent systems. The fourth intelligent system is used to reconstruct and combine multiple feature description contents included in the second information to generate multiple feature description content combinations. The fourth intelligent system is also used to predict the feedback effect of the media information generated by the multiple feature description content combinations respectively.

[0065] Step 3a: Generate second media information based on the fifth intelligent system and the first information. The fifth intelligent system is an intelligent system among at least two second intelligent systems. The fifth intelligent system is used to transform the feature description content corresponding to the first information into second media information that matches the feature description content corresponding to the first information.

[0066] See Figure 4 The third intelligent system can be used to transform unstructured multimodal media information (such as images, videos, and text) into structured, machine-understandable feature descriptions. The second information, corresponding to the feature descriptions, can describe the first media information across multiple dimensions (such as scene, style, characters, and composition). For example, taking the original first media information as the processing object, the third intelligent system deconstructs and analyzes the unstructured first media information along multiple dimensions, filtering out invalid feature information in the analysis results, extracting valid features that reflect the visual characteristics of the first media information, and transforming them into standardized, structured feature descriptions, thereby obtaining the second information.

[0067] The third intelligent system transforms unstructured and non-standardized media information into standardized, quantifiable, and divisible structured feature descriptions, enabling digital analysis of raw media information. This solves the problem that traditional technologies cannot quantify and control the creativity in fragmented media information, transforming ambiguous creative content into editable and optimizable feature descriptions. It effectively addresses the difficulty in accurately controlling and optimizing unstructured information, providing data support for adjusting and optimizing feature descriptions.

[0068] See Figure 4 The fourth intelligent system can receive the second information from the third intelligent system. The fourth intelligent system can recombine the multi-dimensional structured feature descriptions in the second information to form multiple sets of differentiated feature descriptions. Then, by predicting the feedback effects of the media information generated by the combination of multiple feature descriptions, the system can select feature descriptions that ensure the feedback effect of the media information generated based on the first information is no less than that of the media information generated based on the second information, thereby optimizing the feature descriptions according to the feedback effect of the media information.

[0069] The fourth intelligent system can reconstruct the feature description content corresponding to the second information and predict the feedback effect of the media information generated by the reconstructed feature description content. Under the premise of ensuring that the feedback effect of the generated media information is not reduced, an appropriate combination of feature description content is generated as the feature description content corresponding to the first information, so that the generated feature description content corresponding to the first information can produce a better feedback effect, effectively alleviating the problem of insufficient creative diversity in the process of media information generation.

[0070] See Figure 4 The fifth intelligent system can receive the optimized first information from the fourth intelligent system and accurately transform the abstract, structured feature descriptions in the first information into concrete second media information. This ensures that the visual style, copywriting, and layout of the second media information match the corresponding feature descriptions in the first information. By achieving precise transformation from feature descriptions to media information through the fifth intelligent system, the optimized feature description scheme is fully implemented, resulting in significantly differentiated second media information with optimized feedback, effectively alleviating the problem of insufficient creative diversity in the media information generation process.

[0071] In some cases, the second media information can be generated based on the fifth intelligent system and the first information by combining various alternative solutions from one or more of the above situations, including the following steps: generating an image that matches the first information through the fifth intelligent system; generating a video that matches the first information through the fifth intelligent system; and generating a media information title that matches the first information through the fifth intelligent system.

[0072] Optionally, the fifth intelligent system generates an image matching the first information, including: using the fifth intelligent system to call a text-based image model to generate an image matching the first information. The fifth intelligent system also generates a video matching the first information, including: using the fifth intelligent system to call a video generation interface to generate a video matching the first information, wherein the video generation interface is connected to an intelligent system for video generation or video synthesis.

[0073] The text-based image generation model is a pre-trained multimodal generative model with built-in image generation algorithms and feature matching logic. It can receive structured feature descriptions or text commands as input and generate static images that match the input feature descriptions. The text-based image generation model is adaptable to the generation needs of various image-based media information such as posters, advertising images, and product display images. The Fifth Intelligence System directly triggers the text-based image generation model to run through built-in commands.

[0074] In some scenarios, the fifth intelligent system can first parse the initial information, extracting structured feature descriptions related to image generation, including but not limited to style, elements, color, layout, and text. It then calls upon the built-in text-to-image model to issue media information generation instructions, converting the feature descriptions corresponding to the initial information into input parameters recognizable by the text-to-image model. Upon receiving the instructions, the text-to-image model initiates the generation process, outputting image-based media information that matches the feature descriptions corresponding to the initial information, thus completing the generation of the second media information. By calling the text-to-image model through the fifth intelligent system, accurate conversion from structured features to static images is achieved. The generated second media information matches the optimized initial information, ensuring optimized feedback of the second media information. Simultaneously, no manual intervention is required, achieving automated generation of image-based media information, adapting to the needs of generating large batches of advertising images and posters.

[0075] The video generation interface is a standardized data transmission and invocation interface used to connect the fifth intelligent system with external intelligent systems used for video generation models and video compositing. The intelligent system used for video generation or compositing can be an intelligent system with video generation, video editing and compositing, and dynamic material rendering functions. By receiving instructions from the fifth intelligent system through the video generation interface, and based on the feature description content corresponding to the first information, it can invoke the intelligent system used for video generation or compositing to generate dynamic advertising videos, short video creative materials, and composite promotional videos.

[0076] In some scenarios, the fifth intelligent system extracts video-related feature indicators from the first information, which may include, but are not limited to, duration, visual style, frame content, and background music requirements. The feature descriptions corresponding to the first information extracted through the video generation interface are transmitted to the connected intelligent system for video generation or video compositing. The intelligent system for video generation or video compositing completes the video production and then sends the generated video media information back to the fifth intelligent system through the video generation interface, forming the second media information. By connecting to external intelligent systems through the video generation interface, the generation format is expanded, enabling dynamic production of video media information. This facilitates flexible replacement of the intelligent system for video generation, resulting in a highly scalable architecture. The entire process is automated, requiring no manual debugging. While ensuring the media information matches the feature descriptions corresponding to the first information, it improves the efficiency of media information generation, balancing the diversity of media information with process controllability.

[0077] In addition to the aforementioned technical effects, the technical solution in this scenario also achieves the following: by independently encapsulating various generation capabilities related to media information generation into corresponding intelligent systems, and relying on the first intelligent system to achieve unified global orchestration and collaborative management, a fully automated multimodal media information generation process is constructed. Compared to the traditional monolithic model generation architecture and fragmented manual operation mode, this solution not only decouples the single intelligent system from the overall architecture, providing high flexibility and scalability, but also allows for the direct addition, deletion, or replacement of corresponding functional intelligent systems in response to changing needs. For example, it can quickly integrate with intelligent systems used for video generation without reconstructing the overall framework, enabling rapid adaptation to the generation needs of multimodal media information in different scenarios and formats. Furthermore, by unifying the operation of each second intelligent system through the first intelligent system, the entire process of media information generation can be automated without human intervention, effectively avoiding human operation errors and process continuity gaps, significantly shortening the multimodal media information generation cycle, and significantly improving the efficiency and stability of media information generation.

[0078] In some cases, it can be combined with the various alternatives in one or more of the above situations. At least one intelligent system may include a first model. The first model is obtained by reinforcement learning training with third information as input and fourth information as output. The first model is a machine learning model. The first model is used to guide at least one intelligent system to adjust the feature description content corresponding to the second information. The third information is obtained by extracting structured feature description content from the third media information based on multiple dimensions. The fourth information is used to indicate the reward score determined based on the feedback effect of the third media information associated with the third information. The third media information is the media information that has been used and has produced feedback effect before triggering the media information generation task request.

[0079] See Figure 5 The third information is represented in explicit, structured, and interpretable natural language. The third information can be input data for reinforcement learning training of the first model. It includes structured feature descriptions, which can be feature descriptions from multiple dimensions such as style, composition, text, color, and layout. The corresponding feature descriptions for the third information can be obtained by extracting multi-dimensional structured features from already used third-party media information. The corresponding feature descriptions for the third information possess the characteristics of being quantifiable, parsable, and input into the first model. The third-party media information can be finished media information that has already been used online or offline and has generated clear and quantifiable feedback data before the current media information generation task request is triggered, possessing the characteristic of traceable real-world feedback effects.

[0080] See Figure 5 The fourth piece of information can be the output of reinforcement learning training on the first model. It is a reward score quantified based on the feedback effect of the third media information, representing a reward signal in reinforcement learning training. A better feedback effect from the third media information results in a higher reward score, and vice versa. The feedback effect of the third media information is used to positively guide the first model in optimizing its parameters. The feedback effect of the third media information can be a comprehensive evaluation value calculated using multiple posterior metrics such as impressions, click-through rate, and conversion rate. In this way, the first model learns to "evaluate" which combinations of feature descriptions are more likely to generate good feedback effects for the media information.

[0081] See Figure 5 The system selects third-party media information that has been used and has provided feedback. It extracts structured feature descriptions from this third-party media information to obtain third-party information, which serves as the input to the reward model. The reward score, quantified based on the feedback effect of the third-party media information, is used as the output of the reward model. Through multiple rounds of iterative training, the parameters of the reward model are optimized, allowing the model to learn the optimization rules based on feedback effects. After training, the trained first model is embedded into at least one intelligent system. When adjusting the feature descriptions corresponding to the second-party information, the first model outputs optimization guidance, directing at least one intelligent system to adjust the feature descriptions corresponding to the second-party information, thus achieving feedback effect-oriented feature optimization. The first model can be configured in the fourth intelligent system within at least one intelligent system, where the first model outputs optimization guidance to direct the fourth intelligent system to adjust the feature descriptions corresponding to the second-party information.

[0082] See Figure 5The fourth intelligent system may include a feature description content generation model, which is a policy network capable of generating new combinations of feature description content. During training, the feature description content generation model produces new combinations of feature description content. These new combinations are then fed into a pre-trained reward model for scoring. The fourth intelligent system uses reinforcement learning algorithms, employing the reward scores returned by the reward model as positive or negative incentives to update the parameters of the feature description content generation model. Essentially, this process allows the feature description content generation model to continuously explore and learn how to generate feature description content combinations that yield higher expected rewards.

[0083] The technical effect of the above method is that by training a reward model to evaluate the feedback effect of the media information generated by the combination of feature description content, and using the reward model to guide a feature description content generation model to update its parameters through algorithms such as policy gradient, the feature description content generation model tends to generate feature description content combinations with high expected rewards, thereby ensuring that the media information generated by the feature description content corresponding to the first information after the feature description content corresponding to the second information is adjusted according to the fourth intelligent system has a better feedback effect.

[0084] In some cases, it can be combined with various alternative solutions from one or more of the above scenarios. At least one intelligent system may include a second model. The second model is a model built with third information as input and fourth information as output. The second model is used to indicate the mapping relationship between multiple pieces of third information and the fourth information associated with the third information. The second model is used to guide at least one intelligent system to adjust the feature description content corresponding to the second information. The third information is obtained by extracting structured feature description content from third media information based on multiple dimensions. The fourth information is used to indicate the reward score determined based on the feedback effect of the third media information associated with the third information. The third media information is media information that has been used and has produced feedback effects before triggering the media information generation task request. The second model can be configured in the fourth intelligent system of at least one intelligent system. The second model outputs optimization guidance to guide the fourth intelligent system to adjust the feature description content corresponding to the second information. Optionally, the second model is a multiple regression model built with third information as input and fourth information as output.

[0085] In some cases, it can be combined with various alternatives from one or more of the above scenarios. At least one intelligent system can be configured with a multi-armed slot machine algorithm, treating each or each combination of feature descriptions as an "arm." Through online exploration and utilization, the multi-armed slot machine algorithm dynamically allocates more generation opportunities to well-performing feature description combinations, thereby guiding at least one intelligent system to adjust the feature descriptions corresponding to the second information. The multi-armed slot machine algorithm can be configured in a fourth intelligent system within at least one intelligent system, where the algorithm outputs optimization guidance to direct the fourth intelligent system to adjust the feature descriptions corresponding to the second information.

[0086] In some cases, it may be combined with various alternatives in one or more of the above situations, and at least one intelligent system includes a first model and / or a second model.

[0087] In some cases, it can be combined with the various alternatives in one or more of the above cases, see [reference]. Figure 5 The media information generation method provided in this article may also include the following steps: Obtain first identification information, which is included in the media information generation task request. The first identification information is related to the usage scenario type of the media information requested to be generated by the media information generation task request. Based on the first identification information, determine at least one first model and / or second model adopted by the intelligent system. The first model and / or second model associated with different first identification information are different.

[0088] The technical advantage of the above approach is that it allows for the independent training and invocation of reward models with varying preferences for different usage scenarios of media information, achieving precise binding between different reward models and specific usage scenario types. By relying on the quantitative evaluation and guiding role of the reward models corresponding to usage scenario types, at least one intelligent system can be precisely guided to adjust the feature description content corresponding to the second information. This fully leverages the preference differences among different usage scenario types, enhances the relevance and attractiveness of the feature description content corresponding to the first information, strengthens the matching degree between the feature description content corresponding to the first information and the usage scenario type, and thus effectively optimizes the conversion efficiency from the feature description content corresponding to the first information to media information.

[0089] In some cases, it can be combined with the various alternatives in one or more of the above cases, see [reference]. Figure 5 The media information generation method provided in this article may also include the following steps: Acquire multiple third-media information and extract third information from the third-media information; and quantify the feedback effect after using the third-media information into a reward score to obtain fourth information, and use the third information and fourth information to train the first model.

[0090] The technical effect of the above method is that it automatically collects historical media information feedback data, transforms scattered feedback data into reward signals, and eliminates the need for manual annotation and intervention. By training the reward model based on the media information feedback data, it can accurately learn the correlation between feature description content and feedback effect, allowing the feature description content generation strategy to be continuously iterated and optimized, and continuously improving the quality of feature description content and the conversion of feedback effect.

[0091] In some cases, it can be combined with the various alternatives in one or more of the above cases, see [reference]. Figure 5 At least one intelligent system also includes a sixth intelligent system, which, after generating the second media information based on the first media information, may further include the following steps: The sixth intelligent system filters out fourth-generation media information that does not meet preset requirements, and stores fifth-generation media information that meets preset requirements in the media information database. Preset requirements include media information specification requirements and media information quality requirements.

[0092] Optionally, filtering out fourth media information that does not meet preset requirements and storing fifth media information that meets preset requirements in a media information database includes: filtering out fourth media information that does not meet preset media information specifications and / or has media information quality defects, and storing fifth media information that meets preset media information specifications and / or does not have media information quality defects in a media information database for later use.

[0093] The technical advantages of the above approach are as follows: by setting up an independent sixth intelligent system as an automated quality inspection node, an automated content security inspection barrier is built after the feature description content is generated and before it is put into storage. Relying on the detection capabilities of the sixth intelligent system, compliance and content quality are fully automated, eliminating the need for manual review. It is adapted to automated media information generation scenarios and can accurately intercept non-compliant, low-quality, and security-risk problematic media information, avoiding usage risks, effectively ensuring content security, and preventing negative impacts from inferior content. At the same time, it forms a complete quality inspection closed loop of generation-quality inspection-distribution-storage, significantly improving the controllability and stability of the automated production process and reducing quality inspection costs and the probability of oversight.

[0094] Figure 6 A schematic diagram of a media information generation device is provided. This device is applicable to media information generation scenarios that require both large-scale and differentiated generation of media information, and where the generated media information needs to consider both feedback effects and creative diversity. The device can be implemented in software and / or hardware, optionally through electronic devices such as mobile terminals, PCs, or servers. Figure 6 As shown, the media information generation device in this article may specifically include the following: The acquisition module 610 is used to determine the first media information in response to the media information generation task request; The generation module 620 is used to generate second media information based on the first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first target. The second information is determined based on structured feature description content extracted from the first media information from multiple dimensions. The first target indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0095] In some cases, it can be combined with various alternatives in one or more of the above situations. The at least one intelligent system includes a first intelligent system and at least two second intelligent systems. The generation of second media information based on the first media information includes: Based on the first intelligent system, a first task objective and a first execution order are determined. The first execution order is associated with the first task objective. The first task objective indicates the task objective obtained from parsing the task request generated from the media information. The first execution order is used to call at least two second intelligent systems and instruct the at least two second intelligent systems to collaboratively execute their respective associated tasks. The first intelligent system is also used to detect the execution status of the associated tasks of the at least two second intelligent systems. According to the first execution order, second media information is generated based on the first media information and the at least two second intelligent systems.

[0096] In some cases, it can be combined with various alternatives in one or more of the above situations to generate second media information based on the first media information and the at least two second intelligent systems, including: The second information is generated based on the third intelligent system and the first media information. The third intelligent system is an intelligent system among the at least two second intelligent systems. The third intelligent system is used to transform unstructured media information into structured feature description content. First information is generated based on the fourth intelligent system and the second information. The fourth intelligent system is an intelligent system among the at least two second intelligent systems. The fourth intelligent system is used to reconstruct and combine multiple feature description contents included in the second information to generate multiple feature description content combinations. The fourth intelligent system is also used to predict the feedback effect of the media information generated by the multiple feature description content combinations respectively. The fifth intelligent system generates second media information based on the first information. The fifth intelligent system is an intelligent system among the at least two second intelligent systems. The fifth intelligent system is used to convert the feature description content corresponding to the first information into second media information that matches the feature description content corresponding to the first information.

[0097] In some cases, it can be combined with one or more of the alternative solutions mentioned above to generate second media information based on the fifth intelligent system and the first information, including: The fifth intelligent system generates an image that matches the first information. The fifth intelligent system generates a video that matches the first information.

[0098] In some cases, it may be combined with various alternatives in one or more of the above situations, and the at least one intelligent system includes a first model and / or a second model; The first model is obtained by reinforcement learning with the third information as input and the fourth information as output. The first model is used to guide the at least one intelligent system to adjust the feature description content corresponding to the second information. The third information is obtained by extracting structured feature description content of the third media information based on multiple dimensions. The fourth information is used to indicate the reward score determined based on the feedback effect of the third media information associated with the third information. The third media information is media information that has been used and has produced feedback effect before the media information generation task request is triggered. The second model is a model built with third information as input and fourth information as output. The second model is used to indicate the mapping relationship between multiple pieces of third information and the fourth information associated with the third information. The second model is used to guide the at least one intelligent system to adjust the feature description content corresponding to the second information.

[0099] In some cases, the device may be combined with one or more of the alternative solutions described above, and the device further includes: Obtain first identification information, which is included in the media information generation task request, and the first identification information is related to the usage scenario type of the media information requested to be generated by the media information generation task request; Based on the first identification information, the first model and / or second model adopted by the at least one intelligent system are determined, and the first model and / or second model associated with different first identification information are different.

[0100] In some cases, the device may be combined with one or more of the alternative solutions described above, and the device further includes: Acquire multiple third-party media information and extract third-party information from the third-party media information; The feedback effect after using the third media information is quantified into a reward score to obtain the fourth information, and the third information and the fourth information are used to train the first model.

[0101] In some cases, the first information and the second information can be combined with the alternative schemes in one or more of the above situations, and the first information and the second information can be represented in natural language form; or the first information and the second information can be represented in vector feature form.

[0102] In some cases, it can be combined with various alternatives in one or more of the above situations. The at least one intelligent system is a third model, which includes a machine learning model for receiving, processing, and / or generating at least two different modalities of information. The generation of second media information based on the first media information includes: Obtain first prompt information, which includes third information, fifth information, first task objective, and sixth information. The first prompt information is associated with the first task objective, which indicates the task objective obtained from parsing the media information generation task request. The third information is obtained by extracting structured feature descriptions of the third media information based on multiple dimensions. The fifth information is used to indicate the feedback effect of the third media information associated with the third information. The sixth information is used to indicate the adjustment and optimization strategy for the feature descriptions corresponding to the second information. The third media information is media information that has been used and has generated feedback effects before the media information generation task request is triggered. The second media information is generated based on the first prompt information, the third model, and the first media information.

[0103] In some cases, it can be combined with various alternative solutions in one or more of the above situations. The at least one intelligent system also includes a sixth intelligent system. After generating the second media information based on the first media information, it further includes: filtering the fourth media information that does not meet the preset requirements through the sixth intelligent system, and storing the fifth media information that meets the preset requirements into the media information database.

[0104] In the above method, structured feature extraction of the first media information from multiple dimensions yields the second information. This method clearly identifies the feature descriptions of different dimensions within the first media information. Then, at least one intelligent system adjusts and optimizes the feature descriptions corresponding to the second information to obtain the first information, with the first objective indicating that the feedback effect of the media information generated based on the first information is no less than that generated based on the second information. Since the feature descriptions corresponding to the second information are multi-dimensional, discrete, and independently adjustable, at least one intelligent system can flexibly adjust and optimize them while maintaining the feedback effect, allowing the generated feature descriptions corresponding to the first information to produce better feedback. This effectively alleviates the problems of homogenization and insufficient creative diversity in the media information generation process. Because the feature descriptions corresponding to the first information are structured feature descriptions optimized based on feedback effects, the second media information generated by at least one intelligent system based on these feature descriptions exhibits differences in presentation, visual style, and copywriting. This addresses the problems of singular creative forms and insufficient creative diversity in the media information generation process while ensuring the feedback effect of the generated second media information is not reduced as much as possible.

[0105] The aforementioned media information generation apparatus can execute the media information generation method provided in any embodiment of this document, and has the corresponding functional modules and beneficial effects for executing the media information generation method.

[0106] It is worth noting that the various units and modules included in the above-mentioned interactive device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments in this article.

[0107] The following is for reference. Figure 7 This document illustrates a schematic diagram of an electronic device (e.g., a terminal device or server) 700 suitable for implementing the aforementioned media information generation method. The terminal device referred to herein may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments described herein.

[0108] like Figure 7As shown, the electronic device 700 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage device 708 into a random access memory (RAM) 703. The RAM 703 also stores various programs and data required for the operation of the electronic device 700. The processing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0109] Typically, the following devices can be connected to I / O interface 705: input devices 706 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 707 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 708 including, for example, magnetic tapes, hard disks, etc.; and communication devices 709. Communication device 709 allows electronic device 700 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 7 An electronic device 700 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0110] In particular, according to embodiments of this document, the processes described in the above-referenced flowcharts can be implemented as computer software programs. For example, the technical solutions of this document include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the media information generation method shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 709, or installed from a storage device 708, or installed from a ROM 702. When the computer program is executed by the processing device 701, it performs the functions defined in the media information generation method of the embodiments of this document.

[0111] The names of messages or information exchanged between multiple devices in this document are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0112] The electronic device provided in this embodiment and the media information generation method provided in the above technical solutions belong to the same inventive concept. Technical details not described in detail in this document can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.

[0113] This article provides a computer storage medium on which a computer program is stored, which, when executed by a processor, implements the media information generation method provided in the above embodiments.

[0114] It should be noted that the computer-readable medium described above can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this document, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0115] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0116] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0117] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: respond to a media information generation task request, determine first media information; generate second media information based on the first media information, wherein the second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information, the first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system, and the at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective, the second information is determined based on structured feature description content extracted from the first media information from multiple dimensions, and the first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

[0118] Computer program code for performing the operations described herein may be written in one or more programming languages ​​or a combination thereof, including but not limited to object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer may be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this document. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0120] The modules or units described herein can be implemented in software or hardware. The names of modules or units do not necessarily constitute a limitation on the module or unit itself.

[0121] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0122] In the context of this document, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0123] The above description is merely a preferred embodiment and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure herein is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed herein that have similar functions.

[0124] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this document. Certain features described in the context of individual implementations may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented individually or in any suitable sub-combination in multiple implementations.

[0125] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for generating media information, the method comprising: In response to the media information generation task request, determine the first media information; Second media information is generated based on the first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined based on structured feature description content extracted from the first media information from multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

2. The method according to claim 1, wherein the at least one intelligent system comprises a first intelligent system and at least two second intelligent systems, and the step of generating second media information based on the first media information comprises: Based on the first intelligent system, a first task objective and a first execution order are determined. The first execution order is associated with the first task objective. The first task objective indicates the task objective obtained from parsing the task request generated from the media information. The first execution order is used to call at least two second intelligent systems and instruct the at least two second intelligent systems to collaboratively execute their respective associated tasks. The first intelligent system is also used to detect the execution status of the associated tasks of the at least two second intelligent systems. According to the first execution order, second media information is generated based on the first media information and the at least two second intelligent systems.

3. The method according to claim 2, wherein generating second media information based on the first media information and the at least two second intelligent systems comprises: The second information is generated based on the third intelligent system and the first media information. The third intelligent system is an intelligent system among the at least two second intelligent systems. The third intelligent system is used to transform unstructured media information into structured feature description content. First information is generated based on the fourth intelligent system and the second information. The fourth intelligent system is an intelligent system among the at least two second intelligent systems. The fourth intelligent system is used to reconstruct and combine multiple feature description contents included in the second information to generate multiple feature description content combinations. The fourth intelligent system is also used to predict the feedback effect of the media information generated by the multiple feature description content combinations respectively. The fifth intelligent system generates second media information based on the first information. The fifth intelligent system is an intelligent system among the at least two second intelligent systems. The fifth intelligent system is used to convert the feature description content corresponding to the first information into second media information that matches the feature description content corresponding to the first information.

4. The method according to claim 3, wherein generating second media information based on the fifth intelligent system and the first information comprises: The fifth intelligent system generates an image that matches the first information. The fifth intelligent system generates a video that matches the first information.

5. The method according to any one of claims 1-4, wherein the at least one intelligent system comprises a first model and / or a second model; The first model is obtained by reinforcement learning with the third information as input and the fourth information as output. The first model is used to guide the at least one intelligent system to adjust the feature description content corresponding to the second information. The third information is obtained by extracting structured feature description content of the third media information based on multiple dimensions. The fourth information is used to indicate the reward score determined based on the feedback effect of the third media information associated with the third information. The third media information is media information that has been used and has produced feedback effect before the media information generation task request is triggered. The second model is a model built with third information as input and fourth information as output. The second model is used to indicate the mapping relationship between multiple pieces of third information and the fourth information associated with the third information. The second model is used to guide the at least one intelligent system to adjust the feature description content corresponding to the second information.

6. The method according to claim 5, further comprising: Obtain first identification information, which is included in the media information generation task request, and the first identification information is related to the usage scenario type of the media information requested to be generated by the media information generation task request; Based on the first identification information, the first model and / or second model adopted by the at least one intelligent system are determined, and the first model and / or second model associated with different first identification information are different.

7. The method according to claim 5, further comprising: Acquire multiple third-party media information and extract third-party information from the third-party media information; The feedback effect after using the third media information is quantified into a reward score to obtain the fourth information, and the third information and the fourth information are used to train the first model.

8. The method according to claim 1, wherein the first information and the second information are represented in natural language form; or, the first information and the second information are represented in vector feature form.

9. The method according to claim 1, wherein the at least one intelligent system is a third model, the third model comprising a machine learning model for receiving, processing, and / or generating at least two different modalities of information, and the generation of second media information based on the first media information comprising: Obtain first prompt information, which includes third information, fifth information, first task objective, and sixth information. The first prompt information is associated with the first task objective, which indicates the task objective obtained from parsing the media information generation task request. The third information is obtained by extracting structured feature descriptions of the third media information based on multiple dimensions. The fifth information is used to indicate the feedback effect of the third media information associated with the third information. The sixth information is used to indicate the adjustment and optimization strategy for the feature descriptions corresponding to the second information. The third media information is media information that has been used and has generated feedback effects before the media information generation task request is triggered. The second media information is generated based on the first prompt information, the third model, and the first media information.

10. The method according to claim 1, wherein the at least one intelligent system further comprises a sixth intelligent system, and after generating the second media information based on the first media information, further comprises: The sixth intelligent system filters out fourth media information that does not meet the preset requirements and stores fifth media information that meets the preset requirements in the media information database.

11. A media information generation apparatus, the apparatus comprising: The acquisition module is used to determine the first media information in response to the media information generation task request; A generation module is used to generate second media information based on the first media information. The second media information is media information generated by at least one intelligent system based on the feature description content corresponding to the first information. The first information is obtained by adjusting the feature description content corresponding to the second information through the at least one intelligent system. The at least one intelligent system adjusts the feature description content corresponding to the second information according to a first objective. The second information is determined based on structured feature description content extracted from the first media information from multiple dimensions. The first objective indicates that the feedback effect of the media information generated based on the first information is not lower than the feedback effect of the media information generated based on the second information.

12. An electronic device, the electronic device comprising: At least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-10.

13. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the method of any one of claims 1-10.

14. A computer program product comprising a computer program that, when executed by a processor, implements the method as described in any one of claims 1-10.