Multimedia resource generation method and apparatus, electronic device, and storage medium

CN115482324BActive Publication Date: 2026-08-18BAIDU USA LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211207353.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-30
Publication Date
2026-08-18
Estimated Expiration
2042-09-30

AI Technical Summary

Technical Problem

[0003]然而,当前基于AI网络模型生成的图片稳定性差,且仅能生成少量的图片,生成的视频效果差

Benefits of technology

[0019] In this disclosure, the scene to which the target text is applicable is located based on multiple grid blocks of the target scene, a suitable 3D scene description file is generated, and then a 3D game engine is used for rendering, which can generate multimedia resources with stable effects and reliable image quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115482324B_ABST
    Figure CN115482324B_ABST
Patent Text Reader

Abstract

The present disclosure provides a multimedia resource generation method and device, electronic equipment and storage medium, relates to the technical field of artificial intelligence, in particular to the technical field of deep learning, natural language processing, intelligent search and the like. The specific implementation scheme is: performing matching operation on the target text and semantic information of a plurality of grid blocks in a target scene to obtain a target grid block matched with the target text; wherein each grid block is a part of a scene area of the target scene; determining a three-dimensional scene description file of the target grid block based on feature information about a scene element in the target text; and generating a multimedia resource based on the three-dimensional scene description file of the target grid block and a three-dimensional game engine. In the present disclosure, the target text is positioned based on a plurality of grid blocks of a target scene to determine a scene to which the target text is applicable, a suitable three-dimensional scene description file is generated, and then a three-dimensional game engine is used for rendering, so that a multimedia resource with stable effect and reliable picture quality can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning, natural language processing, and intelligent search. Background Technology

[0002] Given a text, AI (Artificial Intelligence) network models can be used to generate an image corresponding to that text. For example, if the text describes "a girl dancing," the AI ​​network model can generate an image of a girl dancing.

[0003] However, current AI network models generate images with poor stability and can only produce a limited number of images, resulting in poor video quality. Therefore, how to generate multimedia resources based on given text remains to be studied. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for generating multimedia resources.

[0005] According to one aspect of this disclosure, a multimedia resource generation method is provided, comprising:

[0006] The semantic information of the target text is matched with that of multiple grid blocks in the target scene to obtain the target grid blocks that match the target text; where each grid block is a part of the scene region of the target scene.

[0007] Based on the feature information about scene elements in the target text, determine the 3D scene description file of the target mesh block;

[0008] Multimedia resources are generated based on the 3D scene description file of the target mesh and the 3D game engine.

[0009] According to another aspect of this disclosure, a multimedia resource generation apparatus is provided, comprising:

[0010] The matching module is used to match the semantic information of the target text with that of multiple grid blocks in the target scene to obtain the target grid blocks that match the target text; where each grid block is a part of the scene region of the target scene.

[0011] The determination module is used to determine the 3D scene description file of the target mesh block based on the feature information about scene elements in the target text;

[0012] The generation module is used to generate multimedia resources based on the target mesh block-based 3D scene description file and 3D game engine.

[0013] According to another aspect of this disclosure, an electronic device is provided, comprising:

[0014] At least one processor; and

[0015] The memory is communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the multimedia resource generation method of any embodiment of the present disclosure.

[0017] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause the computer to perform a multimedia resource generation method according to any embodiment of this disclosure.

[0018] According to another aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a multimedia resource generation method according to any embodiment of this disclosure.

[0019] In this disclosure, the scene to which the target text is applicable is located based on multiple grid blocks of the target scene, a suitable 3D scene description file is generated, and then a 3D game engine is used for rendering, which can generate multimedia resources with stable effects and reliable image quality.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0021] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein:

[0022] Figure 1 This is a schematic flowchart of a multimedia resource generation method according to an embodiment of the present disclosure;

[0023] Figure 2 This is a schematic flowchart of a multimedia resource generation method according to an embodiment of the present disclosure;

[0024] Figure 3 This is a schematic diagram of constructing a multi-level grid of target blocks according to an embodiment of the present disclosure;

[0025] Figure 4 This is a schematic flowchart of a multimedia resource generation method according to an embodiment of the present disclosure;

[0026] Figure 5 This is a flowchart of a multimedia resource generation method according to an embodiment of the present disclosure;

[0027] Figure 6 This is a schematic diagram of the structure of a multimedia resource generation apparatus according to an embodiment of the present disclosure;

[0028] Figure 7 This is another schematic diagram of a multimedia resource generation apparatus according to an embodiment of the present disclosure;

[0029] Figure 8 This is a block diagram of an electronic device used to implement the multimedia resource generation method of the embodiments of this disclosure. Detailed Implementation

[0030] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0031] The multimedia resources involved in this disclosure may include images or videos. To enable the generation of stable-quality multimedia resources based on given text, this disclosure provides a multimedia resource generation method.

[0032] This method determines the required 3D scene for a given text from a known target scene and uses a game engine to generate multimedia resources. For example... Figure 1 The diagram shown illustrates the process of this method, including:

[0033] S101, perform a matching operation between the target text and the semantic information of multiple grid blocks in the target scene to obtain the target grid block that matches the target text; wherein, each grid block is a part of the scene area of ​​the target scene.

[0034] The target text is the given text, and there are several ways to obtain it. For example, you can obtain a link and then identify the target text from the page corresponding to that link. Another example is to capture a speech signal and then convert it into target text. Yet another example is to obtain a text file and extract its text information to obtain the target text.

[0035] Of course, the method of obtaining the target text is not limited in the embodiments disclosed herein.

[0036] In this embodiment, the target scene can be a large scene, such as a city, which may include sub-scenes like theaters, concert halls, schools, and studios. To accurately determine the scene to which the target text applies, this embodiment divides the target scene into different grid blocks to create different sub-scenes. Each grid block represents a portion of the target scene, and each grid block has corresponding semantic information used for matching with the target text. The grid block that best matches the target text is then selected as the target grid block.

[0037] During matching, the semantic matching degree between the target text and the grid blocks can be compared. In practice, a deep learning-based Natural Language Processing (NLP) model can be used to determine the matching degree between the target text and the semantic information. For example, an NLP model can be used to convert the target file into feature vectors and the semantic information into feature vectors, and the matching degree between the two feature vectors can be determined by comparing their cosine similarity.

[0038] S102, Based on the feature information about scene elements in the target text, determine the 3D scene description file of the target mesh block.

[0039] The 3D scene description file includes scene elements, as well as the position and appearance information of each scene element within the scene. This appearance information includes, for example, the size, shape, and color of the scene elements. For instance, in a theater scene, the appearance information of the curtain could include its color and shape.

[0040] Therefore, it can be understood that the 3D scene description file is used to describe the layout and appearance of different elements within the 3D scene required by the target text using a structured language.

[0041] S103 generates multimedia resources based on the 3D scene description file and 3D game engine of the target mesh blocks.

[0042] The 3D game engine can be, for example, Unreal Engine, and any 3D game engine capable of generating multimedia resources based on a 3D scene description file is applicable to the embodiments of this disclosure.

[0043] In this embodiment, the scene required to generate target text is located within multiple grid blocks of the target scene. The selected 3D scene is a scene from a given target scene resource, not a scene generated by an AI network model. Therefore, the quality of the 3D scene in this embodiment is more guaranteed. Furthermore, since the 3D game engine can render 3D images with high quality, the quality of the generated multimedia resources is guaranteed. Therefore, this disclosure can generate high-quality multimedia resources. Moreover, based on the 3D game engine, videos of arbitrary length can be generated, as well as single or multiple images, so the type of multimedia resources generated is not limited. Thus, the multimedia resource generation method of this embodiment is a general method capable of generating high-quality multimedia resources.

[0044] In this embodiment of the disclosure, a scene resource library, also known as a scene set, can be provided. To facilitate the selection of target scenes suitable for the target text from the scene set, semantic tags can be added to each scene in the scene set. Finding the target scene for the target text can then be implemented as follows:

[0045] Step A1: Extract semantic tags from the target text.

[0046] In practice, the first NLP language model can be used to process the target file to obtain the semantic tags of the target text.

[0047] Step A2 involves matching the semantic tags of the target text with the semantic tags of each scene in the scene set.

[0048] In practice, a second NLP language model can be used to compare the matching degree between the semantic tags of the target text and the semantic tags of the scene. For example, these two tags can be converted into feature vectors, and the matching degree between the target text and the scene can be obtained by comparing the similarity between the feature vectors.

[0049] Step A3: Select the scene that best matches the semantic tags of the target text as the target scene.

[0050] In this embodiment of the disclosure, target scenes suitable for target text can be filtered from the scene set by matching semantic tags. The filtered target scenes semantically satisfy the target text, thereby accurately filtering the target scenes.

[0051] In some embodiments, in order to filter grid blocks suitable for target text from a target scene, the target scene is divided in the following manner: Figure 2 As shown, it includes:

[0052] S201, divide the target scene into multiple grid blocks to obtain the first-level grid blocks;

[0053] S202, perform the following operations repeatedly until the first loop termination condition is met, including:

[0054] S2021 further divides each grid block in the previous level into multiple grid blocks, resulting in...

[0055] The current level of the grid;

[0056] S2022, Determine the grid information of the grid blocks at the current level;

[0057] The grid information includes, for example, at least one of the following: the size of the grid block and the number of elements contained within the grid block.

[0058] S2023, if the grid information at the current level meets the preset requirements, determine that the first loop termination condition is met.

[0059] For example, if the size of the grid block is smaller than a specified size (e.g., the length or width of the grid block is less than 100 meters), the first loop termination condition is determined to be met. Another example is if the number of elements contained within the grid block is less than a preset number, in which case the first loop termination condition is determined to be met. Alternatively, the first loop termination condition can also be determined to be met if both the size of the grid block and the number of elements contained within the grid block are less than a preset number.

[0060] by Figure 3 Let's take an example to illustrate how to divide a grid into blocks. For example... Figure 3 As shown, the target scene is projected onto a 2D plane to obtain a map. The 2D scene is divided into 4 squares, resulting in 4 grid blocks for the first level. Each grid block in the first level is further divided into 4 squares, resulting in 16 grid blocks (4x4) for the second level. This process continues, dividing each grid block in the second level into 4 grid blocks, resulting in 64 grid blocks (16x4) for the third level. This process continues until the smallest grid block obtained is small enough or contains few elements, at which point the division stops; otherwise, it continues, thus dividing the target scene into multiple levels of grid blocks.

[0061] Of course, it should be noted that the number of grid blocks divided into each grid block can be the same or different, and both are applicable to the embodiments of this disclosure.

[0062] By dividing the scene into hierarchical grid blocks, it is easy to match the target grid blocks that match the target text from the target scene using grid blocks as the matching granularity. This enables precise selection of 3D scenes and improves the quality of generated multimedia resources. Furthermore, based on the first loop termination condition, the grid block division hierarchy can be limited to obtain reasonable grid blocks, facilitating the accurate construction of the 3D scene of the target text.

[0063] To facilitate matching with the target text, the semantic information of each grid block can be generated from the elements contained within that grid block. For example, the semantic information of each grid block can be manually labeled. Alternatively, the semantic tags of key elements within the grid block can be extracted as the semantic information of that grid block.

[0064] To improve matching accuracy, in this embodiment of the disclosure, semantic tags of the elements contained within each grid block can be extracted; and, if multiple semantic tags are extracted, the multiple semantic tags are arranged to obtain the semantic information of the grid block. For example, Figure 3 As shown, grid block A1 includes a street, vehicles passing by, a landmark building next to the street, and pedestrians crossing the road next to the building. Therefore, the street, vehicles, buildings, and pedestrians can be considered the semantic information of this grid block. When arranging the labels in order, the semantic labels of multiple elements should be arranged according to the expression of natural language to better describe the grid block. Continuing with... Figure 3 Taking grid block A1 as an example, the semantic tags can be arranged as: street, building, vehicle, pedestrian. Alternatively, a text describing grid block A1 can be generated based on the arrangement of semantic tags. For example, the generated text could be: "There is a street next to the building, there are vehicles on the street, and pedestrians are crossing the road." Thus, the semantic information of the same grid block includes the semantic tags of all elements within that grid block. When matching with the target file, every element within the grid block will participate in the matching, improving matching accuracy and thereby constructing a 3D scene suitable for the target text.

[0065] like Figure 4 As shown, the embodiments of this disclosure can obtain target grid blocks that match the target text based on the following method:

[0066] Perform the following operations in a loop, from the first level to the last level, until the second loop termination condition is met, to obtain the target mesh block:

[0067] S401, when the current level is the first level, perform a matching operation between the target text and the semantic information of the grid blocks of the first level of the target scene to obtain candidate grid blocks that match the target text in the current level.

[0068] S402, when the current level is any level other than the first level, obtain multiple grid blocks in the current level that match the target text in the previous level as a set of candidate grid blocks; and perform a matching operation between the target text and the semantic information of each grid block in the set of candidate grid blocks to obtain the candidate grid blocks that match the target text in the current level.

[0069] S403, if the candidate grid block meets the second loop termination condition, the candidate grid block is determined as the target grid block; the second loop termination condition includes that the candidate grid block of the current level is the grid block of the last level or the matching degree between the candidate grid block of the current level and the target text is greater than the matching degree threshold.

[0070] Continue as Figure 3 As shown, the semantic information of the target text is matched with the four grid blocks of the first level. The target text matches grid block O1. Then it matches the four grid blocks B1-B4 contained in grid block O1. Assuming it matches B3, it matches the four grid blocks A1-A4 of B3. Thus, the target text matches grid block A1 the most closely, and the matching degree satisfies the second loop termination condition. Therefore, grid block A1 is selected as the target grid block.

[0071] In this embodiment of the disclosure, grid blocks are matched layer by layer, and each layer can be limited to the range of grid blocks matched at the previous level, thereby filtering out most grid blocks, which can improve the matching efficiency and quickly match the target grid block.

[0072] After matching the target mesh block, a 3D scene description file needs to be generated in order to accurately generate multimedia resources.

[0073] In one possible implementation, to accurately generate multimedia resources based on the target file, a 3D scene description file (hereinafter referred to as the initial description file) for each mesh block can be stored in a resource library. These initial description files contain key elements of a given scene within the mesh block, the locations of these key elements, and their appearance features. Since the key elements within the given scene may differ slightly from those in the target file, the initial description files can be fine-tuned based on the target file. For example, feature information about scene elements can be extracted from the target text using natural language processing techniques; then, this feature information is used to replace the feature information about scene elements in the 3D scene description text of the target mesh block, resulting in the 3D scene description file.

[0074] For example, in a 3D scene, there is a street with a little girl on it. Suppose the initial description file describes the little girl wearing blue clothes, while the target text describes her wearing orange clothes. Then, the color of the little girl's clothes in the initial description file can be replaced with orange. Therefore, this embodiment of the disclosure can accurately understand the features of the scene elements described in the target text based on NLP technology, thereby enabling the generated 3D scene description file to accurately reproduce the scene described in the target text, thus improving the accuracy of the generated multimedia resources.

[0075] In some embodiments, to stably generate 3D scenes, each grid block can be configured with a preset query, employing a question-and-answer approach to extract feature information of scene elements from the target file. This can be implemented as follows: obtaining the preset query associated with the scene elements of the target grid block; extracting the response results of the preset query from the target text using natural language processing (NLP) technology to obtain the feature information of the scene elements. For example, the target grid block might be a street scene. When the target text is about street interviews, the target scene includes various types of street scenes, such as street scenes of iconic urban areas, street scenes of cultural environments, and street scenes of financial districts. If the target file describes a financial interview, the street scene of financial districts can be located based on the grid block. This scene includes multiple buildings as scene elements, and may also include the interviewer and the interviewee as scene elements. The reporter can be either female or male, so a preset query can be set as "What is the reporter's gender?" Based on this preset query, NLP technology is used to understand the target file and find the answer, thereby determining the reporter's gender. Therefore, the scene elements within the target grid block are all known elements from the resource library. These elements can be pre-set with animated images, resulting in more realistic multimedia resources that better match the description of the target text.

[0076] In summary, based on the preset query, the specific features of the pre-set scene elements can be obtained, thereby accurately understanding the target text's description of the 3D scene and generating high-quality multimedia resources.

[0077] In some possible embodiments, a generative model can also be trained based on the target text, and the generative model can be used to generate a 3D scene description file.

[0078] For example, after obtaining the target text and its corresponding 3D scene description file based on the aforementioned method, a training sample consisting of the target text and the 3D scene description file can be constructed in step B1;

[0079] Step B2: Input the target text into the generation model to obtain the 3D scene description information to be compared predicted by the generation model;

[0080] Step B3: Compare the information of the 3D scene description to be compared with the 3D scene description file to obtain the loss value;

[0081] Step B4: Adjust the generative model based on the loss value. If the generative model meets the training convergence condition, end the training of the generative model.

[0082] In this embodiment, the generative model can be an end-to-end neural network model, which may include an encoder and a decoder. The encoder processes the target text to extract feature information of scene elements, and the decoder decodes this feature information to generate a 3D scene description of the target text to be compared. The parameters of the encoder and decoder can then be determined to optimize the loss.

[0083] In practice, different target scenarios can be trained separately to obtain the corresponding generative models for each target scenario.

[0084] After obtaining the generative model, a generative approach can be used to generate a 3D scene description file based on the target text.

[0085] Therefore, the method provided in this embodiment constructs training samples without manual intervention. Training labels also do not require manual annotation, enabling supervised training of the generative model. This allows the model to accurately generate 3D scene description files, thereby improving the image quality of multimedia resources.

[0086] It should be noted that the multimedia resources in this embodiment are generated based on a 3D game engine. A scene resource library can be constructed using various scene templates provided by the 3D game engine. These scene templates are often realistic and have high image quality, thus ensuring the image quality of the multimedia resources generated by the 3D game engine.

[0087] Furthermore, scene templates in a 3D game engine can be understood as being shared by numerous users, thereby reducing the cost of building a scene resource library.

[0088] Furthermore, 3D game engines can utilize different processor capabilities. When a high-performance 3D game engine is used, multimedia resources can be generated quickly, offering advantages such as faster speed and more stable image quality compared to multimedia resources generated by AI network models.

[0089] For example Figure 5 The diagram shown illustrates the framework for video generation in this embodiment. In this embodiment, for a given text, a target scene can be selected from a basic 3D environment library (i.e., a scene set). Then, the scene is matched with multi-level mesh blocks within the target scene to generate or reconstruct a 3D scene and animation. Finally, elements in the 3D scene are replaced according to the given text to obtain a 3D scene description file. A 3D game engine is used to render the 3D scene description file, thereby generating a video.

[0090] Based on the same technical concept, this disclosure also provides a multimedia resource generation device 600, such as... Figure 6 As shown, it includes:

[0091] The matching module 601 is used to perform a matching operation between the target text and the semantic information of multiple grid blocks in the target scene to obtain the target grid blocks that match the target text; wherein, each grid block is a part of the scene region of the target scene;

[0092] The determination module 602 is used to determine the 3D scene description file of the target mesh block based on the feature information about scene elements in the target text;

[0093] Module 603 is used to generate multimedia resources based on the target mesh block-based 3D scene description file and 3D game engine.

[0094] In some embodiments, Figure 6 On the basis of, such as Figure 7 As shown in the diagram, this disclosure also provides a structural schematic of a multimedia resource generation apparatus 700, which further includes:

[0095] The partitioning module 701 is used to divide the target scene into multiple grid blocks to obtain the first-level grid blocks;

[0096] The first loop module 702 is used to perform the following operations repeatedly until the first loop termination condition is met:

[0097] Each grid block in the previous level is further divided into multiple grid blocks to obtain the grid blocks of the current level;

[0098] Determine the grid information of the grid blocks at the current level;

[0099] If the grid information at the current level meets the preset requirements, the first loop termination condition is determined to be met.

[0100] In some embodiments, the matching module 601 is configured to:

[0101] Perform the following operations in a loop, from the first level to the last level, until the second loop termination condition is met, to obtain the target mesh block:

[0102] When the current level is the first level, the semantic information of the target text is matched with that of the grid blocks in the first level of the target scene to obtain the candidate grid blocks that match the target text in the current level.

[0103] If the current level is any level other than the first level, obtain multiple grid blocks in the current level that match the target text from the previous level as a set of candidate grid blocks; and perform a matching operation between the target text and the semantic information of each grid block in the set of candidate grid blocks to obtain the candidate grid blocks that match the target text in the current level.

[0104] If a candidate grid block satisfies the second loop termination condition, the candidate grid block is determined as the target grid block.

[0105] The second loop termination conditions include either the candidate grid block at the current level being the last grid block at the current level, or the candidate grid block at the current level matching the target text with a matching degree greater than the matching degree threshold.

[0106] In some embodiments, such as Figure 7 As shown, the device also includes a semantic information determination module 703, used for:

[0107] For each grid block, extract the semantic tags of the elements contained within the grid block; and,

[0108] When multiple semantic tags are extracted, they are arranged to obtain the semantic information of the grid block.

[0109] In some embodiments, such as Figure 7 As shown, the device also includes a scene determination module 704, used to determine the target scene based on the following method:

[0110] Extract semantic tags from the target text;

[0111] The semantic tags of the target text are matched with the semantic tags of each scene in the scene set.

[0112] Select the scene that best matches the semantic tags of the target text as the target scene.

[0113] In some embodiments, such as Figure 7 As shown, module 602 includes:

[0114] Natural Language Understanding Unit 705 is used to extract feature information about scene elements in target text based on natural language processing technology;

[0115] The file determination unit 706 is used to replace the feature information about scene elements in the 3D scene description text of the target mesh block with the feature information of scene elements to obtain a 3D scene description file.

[0116] In some embodiments, the natural language understanding unit 705 is used for:

[0117] Retrieve the pre-defined query statements associated with scene elements for the target grid block;

[0118] Natural language processing technology is used to extract the response results of a preset query from the target text, thereby obtaining the feature information of scene elements.

[0119] In some embodiments, such as Figure 7 As shown, the device also includes a training module 707, used for:

[0120] Construct training samples consisting of target text and 3D scene description files;

[0121] Input the target text into the generative model to obtain the 3D scene description information to be compared predicted by the generative model;

[0122] The information of the 3D scene description to be compared is compared with the 3D scene description file to obtain the loss value;

[0123] The generative model is adjusted based on the loss value, and training of the generative model ends when the model meets the training convergence condition.

[0124] The specific functions and examples of each module and unit of the apparatus in this disclosure embodiment can be found in the relevant descriptions of the corresponding steps in the above method embodiments, and will not be repeated here.

[0125] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0126] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0127] like Figure 8 As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.

[0128] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0129] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the method for generating multimedia resources. For example, in some embodiments, the method for generating multimedia resources may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method for generating multimedia resources described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the method for generating multimedia resources by any other suitable means (e.g., by means of firmware).

[0130] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), hybrid programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0131] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0132] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0134] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0135] A computer system may include a client and a server. The client and server are generally located far apart and typically interact via a communication network. The client-server relationship is created by computer programs running on respective computers that have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server incorporating blockchain technology. Embodiments of this disclosure may employ a server to perform a multimedia resource generation method.

[0136] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0137] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for generating multimedia resources, comprising: The semantic information of the target text is matched with that of multiple grid blocks in the target scene to obtain the target grid block that matches the target text. The target scene is divided into multiple levels of grid blocks, and each grid block is a part of the scene area of ​​the target scene. By matching grid blocks level by level, and when the current level is any level other than the first level, each level is limited to the range of grid blocks matched by the previous level, so as to match a candidate grid block that meets the second loop termination condition as the target grid block. The second loop termination condition includes that the candidate grid block of the current level is the grid block of the last level or the matching degree between the candidate grid block of the current level and the target text is greater than the matching degree threshold. Based on the feature information about scene elements in the target text, a 3D scene description file for the target mesh block is determined; Multimedia resources are generated based on the 3D scene description file and 3D game engine of the target mesh blocks.

2. The method according to claim 1, further comprising: The target scene is divided into multiple grid blocks to obtain the first-level grid blocks; Repeat the following operations until the first loop termination condition is met: Each grid block in the previous level is further divided into multiple grid blocks to obtain the grid blocks of the current level; Determine the grid information of the grid blocks at the current level; If the grid information at the current level meets the preset requirements, then the first loop termination condition is determined to be satisfied.

3. The method according to claim 2, wherein, The step of matching the target text with the semantic information of multiple grid blocks in the target scene to obtain the target grid block that matches the target text includes: Perform the following operations in a loop, from the first level to the last level, until the second loop termination condition is met, to obtain the target mesh block: When the current level is the first level, the semantic information of the target text is matched with the grid blocks of the first level of the target scene to obtain candidate grid blocks that match the target text in the current level. If the current level is any level other than the first level, obtain multiple grid blocks in the current level that match the target text in the previous level as a set of candidate grid blocks; and perform a matching operation between the target text and the semantic information of each grid block in the set of candidate grid blocks to obtain candidate grid blocks that match the target text in the current level. If the candidate grid block satisfies the second loop termination condition, the candidate grid block is determined as the target grid block.

4. The method according to claim 2, further comprising: For each grid block, extract the semantic tags of the elements contained within the grid block; and, When multiple semantic tags are extracted, the multiple semantic tags are arranged to obtain the semantic information of the grid block.

5. The method according to any one of claims 1-4, further comprising determining the target scene based on the following method: Extract semantic tags from the target text; The semantic tags of the target text are matched with the semantic tags of each scene in the scene set; The scene that best matches the semantic tags of the target text is selected as the target scene.

6. The method according to claim 1, wherein, The step of determining the 3D scene description file of the target mesh block based on the feature information about scene elements in the target text includes: Feature information about scene elements in the target text is extracted based on natural language processing technology; The feature information of the scene element is used to replace the feature information of the scene element in the 3D scene description text of the target mesh block, thereby obtaining the 3D scene description file.

7. The method according to claim 6, wherein, The extraction of feature information about scene elements from the target text based on natural language processing technology includes: Obtain the preset query statements associated with the scene elements of the target grid block; The response results of the preset query sentence are extracted from the target text using natural language processing technology to obtain the feature information of the scene elements.

8. The method according to claim 1, further comprising: Construct training samples consisting of the target text and the 3D scene description file; The target text is input into the generation model to obtain the 3D scene description information to be compared predicted by the generation model; The information of the three-dimensional scene to be compared is compared with the information of the three-dimensional scene description file to obtain the loss value; The generative model is adjusted based on the loss value, and training of the generative model ends when the generative model meets the training convergence condition.

9. A multimedia resource generation device, comprising: The matching module is used to match the semantic information of target text with multiple grid blocks in the target scene to obtain target grid blocks that match the target text. The target scene is divided into multiple levels of grid blocks, each grid block representing a portion of the target scene. By matching grid blocks level by level, and when the current level is any level other than the first level, each level is limited to the range of grid blocks matched at the previous level, so that a candidate grid block satisfying a second loop termination condition is matched as the target grid block. The second loop termination condition includes either the candidate grid block at the current level being a grid block at the last level or the matching degree between the candidate grid block at the current level and the target text being greater than a matching degree threshold. The determination module is used to determine the three-dimensional scene description file of the target mesh block based on the feature information about scene elements in the target text; The generation module is used to generate multimedia resources based on the 3D scene description file and 3D game engine of the target mesh block.

10. The apparatus according to claim 9, further comprising: A partitioning module is used to divide the target scene into multiple grid blocks to obtain the first-level grid blocks; The first loop module is used to repeatedly perform the following operations until the first loop termination condition is met: Each grid block in the previous level is further divided into multiple grid blocks to obtain the grid blocks of the current level; Determine the grid information of the grid blocks at the current level; If the grid information at the current level meets the preset requirements, then the first loop termination condition is determined to be satisfied.

11. The apparatus according to claim 10, wherein, The matching module is used for: Perform the following operations in a loop, from the first level to the last level, until the second loop termination condition is met, to obtain the target mesh block: When the current level is the first level, the semantic information of the target text is matched with the grid blocks of the first level of the target scene to obtain candidate grid blocks that match the target text in the current level. If the current level is any level other than the first level, obtain multiple grid blocks in the current level that match the target text in the previous level as a set of candidate grid blocks; Then, the target text is matched with the semantic information of each grid block in the candidate grid block set to obtain the candidate grid blocks that match the target text in the current level; If the candidate grid block satisfies the second loop termination condition, the candidate grid block is determined as the target grid block.

12. The apparatus according to claim 10, further comprising a semantic information determination module, configured to: For each grid block, extract the semantic tags of the elements contained within that grid block; and, When multiple semantic tags are extracted, the multiple semantic tags are arranged to obtain the semantic information of the grid block.

13. The apparatus according to any one of claims 9-12, further comprising a scene determination module for determining the target scene based on the following method: Extract semantic tags from the target text; The semantic tags of the target text are matched with the semantic tags of each scene in the scene set; The scene that best matches the semantic tags of the target text is selected as the target scene.

14. The apparatus according to claim 9, wherein, The determining module includes: The natural language understanding unit is used to extract feature information about scene elements from the target text based on natural language processing technology; The file determination unit is used to replace the feature information about the scene element in the 3D scene description text of the target mesh block with the feature information of the scene element to obtain the 3D scene description file.

15. The apparatus according to claim 14, wherein, The natural language understanding unit is used for: Obtain the preset query statements associated with the scene elements of the target grid block; The response results of the preset query sentence are extracted from the target text using natural language processing technology to obtain the feature information of the scene elements.

16. The apparatus of claim 9, further comprising a training module for: Construct training samples consisting of the target text and the 3D scene description file; The target text is input into the generation model to obtain the 3D scene description information to be compared predicted by the generation model; The information of the three-dimensional scene to be compared is compared with the information of the three-dimensional scene description file to obtain the loss value; The generative model is adjusted based on the loss value, and training of the generative model ends when the generative model meets the training convergence condition.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-8.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Method and device for generating dialogue information, equipment and medium

    CN114625855A

  • Method and device for converting text into video, electronic equipment and storage medium

    CN114638232A