Text generation three-dimensional content quality evaluation method, system, medium and terminal

By extracting and fusing the shape, texture and text content consistency features of text-generated three-dimensional content, the problem of the existing technology being unable to effectively evaluate the quality of text-guided three-dimensional content generation is solved, and an accurate assessment of the consistency, authenticity and perceptual quality of three-dimensional content is achieved.

CN119516566BActive Publication Date: 2025-10-21SHANGHAI JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411644056.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-18
Publication Date
2025-10-21
Estimated Expiration
2044-11-18

AI Technical Summary

Technical Problem

Existing quality assessment algorithms cannot effectively evaluate the quality of text-guided generated 3D content, especially under no-reference conditions. It is difficult to evaluate the authenticity of the 3D content, the consistency between the text and the 3D content, and the perceptual quality.

Method used

By obtaining the projection video and text prompt words of the three-dimensional content generated by the text, the pre-trained shape feature extractor, texture feature extractor and text content consistency feature extractor are used to extract shape features, texture features and text content consistency features respectively, and perform feature fusion processing to generate the evaluation score of the three-dimensional content.

Benefits of technology

It achieves effective evaluation of the consistency, authenticity and perceptual quality of text-generated three-dimensional content, and improves the accuracy and efficiency of the evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119516566B_ABST
    Figure CN119516566B_ABST
Patent Text Reader

Abstract

The disclosure provides a quality evaluation method, system, medium and terminal for text-generated three-dimensional content, wherein the method comprises: obtaining a projection video and a text prompt word of the text-generated three-dimensional content; inputting the projection video into a pre-trained shape feature extractor to determine the shape feature of the text-generated three-dimensional content; inputting the projection video into a pre-trained texture feature extractor to determine the texture feature of the text-generated three-dimensional content; inputting the projection video and the text prompt word into a pre-trained text content consistency feature extractor to determine the text content consistency feature of the text-generated three-dimensional content; and performing feature fusion processing on the shape feature, the texture feature and the text content consistency feature to determine the evaluation score of the text-generated three-dimensional content. Through the disclosure, the overall perceptual quality of the text-guided artificial intelligence generated three-dimensional content can be effectively evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a method, system, medium, and terminal for evaluating the quality of three-dimensional content generated by text. Background Art

[0002] Artificial intelligence generated content (AIGC) uses artificial intelligence technology, especially machine learning and deep learning algorithms, to automatically create content such as text, images, audio, video, and three-dimensional objects. For image and video generated content, there are currently a large number of quality evaluation algorithms to evaluate the quality of generated images and videos; at the same time, there are also a relatively rich number of quality evaluation algorithms for traditional three-dimensional content (such as point clouds and meshes). However, there is a severe lack of quality evaluation algorithms for text-guided generated three-dimensional content. Therefore, how to accurately evaluate the quality of these generated three-dimensional contents has become a task that needs to be urgently addressed. Developing a suitable objective quality evaluation method for text-guided generated three-dimensional content can help users screen out three-dimensional content that has high consistency with the prompt text, high authenticity, and good perceived quality, saving users time and costs.

[0003] Compared to the task of evaluating the quality of generated image and video content, the main difference between text-guided 3D content generation and quality evaluation is that the content to be evaluated is 3D content, requiring effective evaluation of the consistency between the text and the generated 3D content, as well as the authenticity and perceptual quality of the generated 3D content. Compared to traditional 3D content quality evaluation tasks, the main differences between text-guided 3D content generation and quality evaluation are that: 3D content is generated using text prompts and has no corresponding reference model, requiring no reference for evaluating the authenticity of the 3D content, the consistency between the text and the 3D content, and the perceptual quality of the 3D content. Furthermore, the generated 3D content is typically expressed implicitly, such as using implicit neural radiation fields, without a fixed format for explicit expression. Therefore, an efficient and effective quality evaluation method is needed.

[0004] Currently, existing quality evaluation algorithms for generated image and video content, as well as quality evaluation algorithms for traditional three-dimensional content, are unable to effectively evaluate the quality of text-guided generated three-dimensional content. Summary of the Invention

[0005] In view of the defects in the prior art, the present invention aims to provide a method, system, medium and terminal for evaluating the quality of three-dimensional content generated by text.

[0006] To achieve the above objectives, according to one aspect of the present disclosure, a method for evaluating the quality of three-dimensional content generated from text is provided, comprising:

[0007] Obtaining projection video and text prompts for text-generated three-dimensional content;

[0008] Inputting the projected video into a pre-trained shape feature extractor to determine shape features of the text to generate three-dimensional content;

[0009] Inputting the projected video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content;

[0010] Inputting the projected video and the text prompt words into a pre-trained text content consistency feature extractor to determine the text content consistency features of the three-dimensional content generated by the text;

[0011] The shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content are subjected to feature fusion processing to determine an evaluation score of the text-generated three-dimensional content.

[0012] Optionally, the evaluation score of the text-generated three-dimensional content includes a consistency quality score, an authenticity quality score, and a perception quality score.

[0013] Optionally, the pre-trained shape feature extractor includes a three-dimensional encoder.

[0014] Optionally, inputting the projection video into a pre-trained shape feature extractor to determine shape features of the text to generate three-dimensional content includes:

[0015] Performing time-domain downsampling processing on the projected video to determine a downsampled video;

[0016] The downsampled video is input into the 3D encoder to determine shape features of the text to generate 3D content.

[0017] Optionally, the pre-trained texture feature extractor includes a first two-dimensional encoder and a second two-dimensional encoder.

[0018] Optionally, inputting the projected video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content includes:

[0019] Determining a front image and a back image of the projected video according to the projected video;

[0020] Inputting the front image into the first two-dimensional encoder to determine the front texture features of the text-generated three-dimensional content;

[0021] Inputting the back side image into the second two-dimensional encoder to determine the back side texture features of the text to generate three-dimensional content;

[0022] The front texture features of the text-generated three-dimensional content and the back texture features of the text-generated three-dimensional content are fused to determine the texture features of the text-generated three-dimensional content.

[0023] Optionally, determining the front image and the back image of the projected video according to the projected video includes:

[0024] Using the first frame image of the projected video as the front image of the projected video;

[0025] The time domain center frame image of the projection video is used as the back image of the projection video.

[0026] Optionally, the pre-trained text content consistency feature extractor includes an image encoder and a text encoder.

[0027] Optionally, inputting the projected video and the text prompt words into a pre-trained text content consistency feature extractor to determine the text content consistency features of the text-generated three-dimensional content includes:

[0028] Using the first frame image of the projected video as the front image of the projected video;

[0029] Inputting the front image of the projected video into the image encoder to determine image features of the text-generated three-dimensional content;

[0030] Inputting the text prompt word into the text encoder to determine text features of the text to generate three-dimensional content;

[0031] The image features of the text-generated three-dimensional content and the text features of the text-generated three-dimensional content are fused to determine the text content consistency features of the text-generated three-dimensional content.

[0032] Optionally, the step of acquiring text to generate a projection video and text prompt words of three-dimensional content includes:

[0033] Rendering the text to generate a projection video of the three-dimensional content using a rendering camera that surrounds the text to generate the three-dimensional content according to a preset projection video resolution size;

[0034] Obtain preset text prompt words for generating the three-dimensional content of the text.

[0035] According to a second aspect of the present disclosure, a quality evaluation system for generating three-dimensional content from text is provided, comprising:

[0036] An acquisition module, used to acquire projection video and text prompt words for generating three-dimensional content from text;

[0037] a shape feature extraction module, configured to input the projection video into a pre-trained shape feature extractor to determine shape features of the text to generate three-dimensional content;

[0038] A texture feature extraction module, configured to input the projection video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content;

[0039] A text content consistency feature extraction module, configured to input the projection video and the text prompt words into a pre-trained text content consistency feature extractor to determine text content consistency features of the text-generated three-dimensional content;

[0040] The quality evaluation module is used to perform feature fusion processing on the shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content to determine the evaluation score of the text-generated three-dimensional content.

[0041] According to a third aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of the method provided in the first aspect of the present disclosure are implemented.

[0042] According to a fourth aspect of the present disclosure, a terminal is provided, including:

[0043] a memory having a computer program stored thereon;

[0044] A processor is used to execute the computer program in the memory to implement the steps of the method provided in the first aspect of the present disclosure.

[0045] Compared with the prior art, the embodiments of the present disclosure have at least one of the following beneficial effects:

[0046] Through the above technical solution, the text-generated three-dimensional content is rendered to obtain a projection video, and the text prompt words corresponding to the generated three-dimensional content are obtained. A pre-trained shape feature extractor is used to obtain the shape features of the text-generated three-dimensional content, a pre-trained texture feature extractor is used to obtain the texture features of the text-generated three-dimensional content, and a pre-trained text content consistency feature extractor is used to obtain the text content consistency of the text-generated three-dimensional content. Finally, the shape features, texture features and text content consistency features of the text-generated three-dimensional content are feature fused to obtain the final evaluation score of the text-generated three-dimensional content, thereby achieving effective evaluation of the consistency, authenticity and perceptual quality of the text-generated three-dimensional content. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Other features, objects and advantages of the present disclosure will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:

[0048] Figure 1 The present invention is a flowchart of a method for evaluating the quality of three-dimensional content generated from text according to an exemplary embodiment.

[0049] Figure 2 The present invention is a process diagram of a method for evaluating the quality of three-dimensional content generated from text according to an exemplary embodiment.

[0050] Figure 3 The figure is a schematic diagram showing a process of extracting shape features by a pre-trained shape feature extractor according to an exemplary embodiment.

[0051] Figure 4 The figure is a schematic diagram of a process of extracting texture features using a pre-trained texture feature extractor according to an exemplary embodiment.

[0052] Figure 5 The present invention is a flowchart illustrating a process of extracting text content consistency features by a pre-trained text content consistency feature extractor according to an exemplary embodiment.

[0053] Figure 6 The present invention is a block diagram of a quality evaluation system for generating three-dimensional content from text according to an exemplary embodiment. DETAILED DESCRIPTION

[0054] The present disclosure is described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art further understand the present disclosure, but are not intended to limit the present disclosure in any way. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the scope of the present disclosure. These modifications and improvements are all within the scope of protection of the present disclosure.

[0055] Figure 1 The present invention is a flowchart of a method for evaluating the quality of three-dimensional content generated from text according to an exemplary embodiment. Figure 2 The present invention is a process diagram of a method for evaluating the quality of three-dimensional content generated from text according to an exemplary embodiment.

[0056] like Figure 1 、 Figure 2 As shown, the present disclosure provides a method for evaluating the quality of three-dimensional content generated by text, including S11 to S15.

[0057] S11, obtaining a projection video and text prompt words for generating three-dimensional content from text.

[0058] The projected video refers to a video rendered by using text to generate three-dimensional content, and the text prompt words refer to preset text content used to generate the three-dimensional content, that is, preset text prompt words.

[0059] S12, inputting the projected video into a pre-trained shape feature extractor to determine the shape features of the text to generate three-dimensional content.

[0060] Among them, the pre-trained shape feature extractor includes a three-dimensional encoder.

[0061] S13, inputting the projected video into a pre-trained texture feature extractor to determine texture features of the text to generate three-dimensional content.

[0062] Among them, the pre-trained texture feature extractor includes a first two-dimensional encoder, a second two-dimensional encoder, a fully connected layer 1, a fully connected layer 2, and an activation function RELU1.

[0063] S14: Input the projected video and text prompt words into a pre-trained text content consistency feature extractor to determine text content consistency features for generating three-dimensional content from the text.

[0064] Among them, the pre-trained text content consistency feature extractor is a multimodal pre-trained neural network, which includes an image encoder, a text encoder, a fully connected layer 1, a fully connected layer 2, and an activation function RELU1.

[0065] S15 , performing feature fusion processing on the shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content to determine an evaluation score of the text-generated three-dimensional content.

[0066] Among them, the evaluation scores of text-generated three-dimensional content include consistency quality score, authenticity quality score and perception quality score.

[0067] The shape features, texture features and text content consistency features of the three-dimensional content generated by the text disclosed in the present invention are fused through the fully connected layer 1, the activation function RELU1, the fully connected layer 2, the activation function RELU2 and the fully connected layer 3 to generate the consistency quality score, authenticity quality score and perception quality score of the three-dimensional content of the text.

[0068] Through the above technical solution, the text-generated three-dimensional content is rendered to obtain a projection video, and the text prompt words corresponding to the generated three-dimensional content are obtained. A pre-trained shape feature extractor is used to obtain the shape features of the text-generated three-dimensional content, a pre-trained texture feature extractor is used to obtain the texture features of the text-generated three-dimensional content, and a pre-trained text content consistency feature extractor is used to obtain the text content consistency of the text-generated three-dimensional content. Finally, the shape features, texture features and text content consistency features of the text-generated three-dimensional content are feature fused to obtain the final evaluation score of the text-generated three-dimensional content, thereby achieving effective evaluation of the consistency, authenticity and perceptual quality of the text-generated three-dimensional content.

[0069] In a possible embodiment, S11 may include S21 to S22.

[0070] S21 , according to a preset projection video resolution size, using a rendering camera that surrounds the text to generate three-dimensional content, rendering the text to generate a projection video of the three-dimensional content.

[0071] The trajectory of the rendering camera is set to surround the text to generate three-dimensional content, and the rendering camera is used to render the text to generate a projection video corresponding to the three-dimensional content.

[0072] The preset projection video resolution size can be set according to the input requirements of the subsequent shape feature extractor, texture feature extractor, and text content consistency feature extractor.

[0073] In the present disclosure, the total number of frames of the rendered projection video is 120 frames to reduce the amount of calculation.

[0074] S22: Obtain preset text prompt words for generating the three-dimensional content of the text.

[0075] Among them, in the present disclosure, the preset text prompt words can be directly obtained.

[0076] Figure 3 The figure is a schematic diagram showing a process of extracting shape features by a pre-trained shape feature extractor according to an exemplary embodiment.

[0077] like Figure 3 As shown, in a possible embodiment, S12 may include S31 to S32.

[0078] S31 , performing time domain downsampling processing on the projected video to determine a downsampled video.

[0079] The time domain downsampling process may adopt a uniform sampling method, and image frames may be selected at intervals in the projected video according to a preset time sequence to form a downsampled video.

[0080] In the present disclosure, the total number of frames of the downsampled video is 12 frames to reduce the amount of calculation.

[0081] S32, inputting the downsampled video into a 3D encoder to determine shape features of the text to generate 3D content.

[0082] In the present disclosure, the 3D encoder may use the Swin3D-S network pre-trained on the KINETICS400 dataset and use bilinear interpolation to downsample the input downsampled video to a resolution of 224×244.

[0083] Thus, the pre-trained shape feature extractor outputs shape features of the text-generated three-dimensional content.

[0084] Figure 4 The figure is a schematic diagram of a process of extracting texture features using a pre-trained texture feature extractor according to an exemplary embodiment.

[0085] like Figure 4 As shown, in a possible embodiment, S13 may include S41 to S44.

[0086] S41, determining a front image and a back image of the projection video according to the projection video.

[0087] As an example, the first frame image of the projection example is used as the front image of the projection video.

[0088] As another example, the time-domain center frame image of the projected video is used as the back image of the projected video.

[0089] S42: Input the front image into a first two-dimensional encoder to determine the front texture features of the text to generate three-dimensional content.

[0090] The first two-dimensional encoder is a pre-trained image neural network, and the first two-dimensional encoder is used to extract front texture features from the front image.

[0091] S43: Input the back image into a second two-dimensional encoder to determine the back texture features of the text to generate three-dimensional content.

[0092] The second two-dimensional encoder is a pre-trained image neural network, and the second two-dimensional encoder is used to extract back texture features from the back image.

[0093] In the present disclosure, both the first two-dimensional encoder and the second two-dimensional encoder use the Swin-S network pre-trained on the IMAGENET1K dataset, and use bilinear interpolation to downsample the input front image and back image to a resolution of 224×244.

[0094] S44 , fusing the front texture features of the text-generated three-dimensional content and the back texture features of the text-generated three-dimensional content to determine the texture features of the text-generated three-dimensional content.

[0095] The fully connected layer 1, activation function RELU1 and fully connected layer 2 are used to fuse the front texture features and the back texture features, and the pre-trained texture feature extractor outputs the texture features of the three-dimensional content of the text.

[0096] Figure 5 The present invention is a flowchart illustrating a process of extracting text content consistency features by a pre-trained text content consistency feature extractor according to an exemplary embodiment.

[0097] like Figure 5As shown, in a possible embodiment, S14 may include S51 to S55.

[0098] S51: Use the first frame image of the projected video as the front image of the projected video.

[0099] S52: Input the front image of the projected video into an image encoder to determine image features of the text to generate three-dimensional content.

[0100] Among them, an image encoder is used to extract the image features of the projected video to extract text from the front image to generate three-dimensional content.

[0101] S53: Input the text prompt words into a text encoder to determine text features for generating three-dimensional content from the text.

[0102] Among them, a text encoder is used to extract text features of three-dimensional content from text prompt words.

[0103] In the present disclosure, the text content consistency feature extractor adopts the CLIP multimodal neural network of weakly supervised image-text data, wherein the image encoder adopts the VisionTransformer architecture, the text encoder adopts the Transformer architecture, and the bilinear interpolation method is used to downsample the image input to the image encoder to a resolution of 224×244.

[0104] S54 , fusing the image features of the text-generated three-dimensional content and the text features of the text-generated three-dimensional content to determine text content consistency features of the text-generated three-dimensional content.

[0105] The fully connected layer 1, activation function RELU1 and fully connected layer 2 are used to fuse the image features and text features of the text-generated three-dimensional content, and the pre-trained text content consistency feature extractor outputs the text content consistency features of the text-generated three-dimensional content.

[0106] In a possible embodiment, S15, the shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content are subjected to feature fusion processing to determine the evaluation score of the text-generated three-dimensional content, which may include: splicing the shape features, texture features, and text content consistency features of the text-generated three-dimensional content, and inputting the spliced ​​features into a neural network composed of a fully connected layer 1, an activation function RELU1, a fully connected layer 2, an activation function RELU2, and a fully connected layer 3, and outputting the consistency quality score, authenticity quality score, and perception quality score of the text-generated three-dimensional content.

[0107] In a possible embodiment, the AIGC-T23DCQA Database is used to verify the effectiveness of a quality evaluation method for generating three-dimensional content from text provided in the present disclosure.

[0108] Among them, the AIGC-T23DCQA Database includes 969 verified 3D contents, which are generated from 170 text prompt words through 6 popular text-to-3D content generation models. These 3D contents are given corresponding subjective quality ratings from the perspectives of perceptual quality, authenticity and text content consistency.

[0109] The evaluation criteria used in the verification test for authenticity, text content consistency and perceived quality include Spearman's Rank-Order Correlation Coefficient (SRCC), Kendall's Rank-Order Correlation Coefficient (KRCC) and Pearson linear correlation coefficient (PLCC).

[0110] The AIGC-T23DCQA database is used to verify traditional image quality assessment algorithms NIQE, ILNIQE, BRISQUE, Resnet-18, Resnet-34, Swin-T, CNNIQA, Swin3D-S, CLIPScore, and the quality assessment method for text-generated three-dimensional content provided in the present invention, namely, the Proposed Method.

[0111] For traditional image quality assessment algorithms, the average score across all frames of the projected video is calculated as the final result. For the remaining quality assessment algorithms requiring training, the AIGC-T23DCQA database is split into training and test sets in a 4:1 ratio. Furthermore, the dataset was randomly split 10 times, and the results were averaged to ensure unbiased performance comparisons.

[0112] The performance test results are shown in the following table:

[0113]

[0114] Table 1

[0115] As can be seen from Table 1, the quality evaluation method for text-generated three-dimensional content proposed in the present disclosure can effectively evaluate the text content consistency, authenticity and perceptual quality of text-generated three-dimensional content.

[0116] Figure 6The present invention is a block diagram of a quality evaluation system for generating three-dimensional content from text according to an exemplary embodiment.

[0117] Based on the same concept, the present disclosure also provides a quality evaluation system 100 for generating three-dimensional content from text, including: an acquisition module 110 , a shape feature extraction module 120 , a texture feature extraction module 130 , a text content consistency feature extraction module 140 , and a quality evaluation module 150 .

[0118] An acquisition module 110 is used to acquire projection video and text prompt words for generating three-dimensional content from text;

[0119] A shape feature extraction module 120 is configured to input the projection video into a pre-trained shape feature extractor to determine shape features of the text-generated three-dimensional content;

[0120] A texture feature extraction module 130 is configured to input the projection video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content;

[0121] A text content consistency feature extraction module 140 is configured to input the projection video and the text prompt words into a pre-trained text content consistency feature extractor to determine text content consistency features of the text-generated three-dimensional content;

[0122] The quality evaluation module 150 is used to perform feature fusion processing on the shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content to determine the evaluation score of the text-generated three-dimensional content.

[0123] Through the above technical solution, the text-generated three-dimensional content is rendered to obtain a projection video, and the text prompt words corresponding to the generated three-dimensional content are obtained. A pre-trained shape feature extractor is used to obtain the shape features of the text-generated three-dimensional content, a pre-trained texture feature extractor is used to obtain the texture features of the text-generated three-dimensional content, and a pre-trained text content consistency feature extractor is used to obtain the text content consistency of the text-generated three-dimensional content. Finally, the shape features, texture features and text content consistency features of the text-generated three-dimensional content are feature fused to obtain the final evaluation score of the text-generated three-dimensional content, thereby achieving effective evaluation of the consistency, authenticity and perceptual quality of the text-generated three-dimensional content.

[0124] Regarding the embodiment of the above system, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0125] Based on the same concept as above, in another embodiment of the present disclosure, a terminal is also provided, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and a method for evaluating the quality of three-dimensional content generated by text when the processor executes the program.

[0126] Optionally, the memory is used to store programs; the memory may include volatile memory (English: volatile memory), such as random-access memory (English: random-access memory, abbreviated: RAM), such as static random-access memory (English: static random-access memory, abbreviated: SRAM), double data rate synchronous dynamic random access memory (English: Double Data Rate Synchronous Dynamic Random Access Memory, abbreviated: DDR SDRAM), etc.; the memory may also include non-volatile memory (English: non-volatile memory), such as flash memory (English: flash memory). The memory is used to store computer programs (such as applications, functional modules, etc. that implement the above-mentioned methods), computer instructions, etc., and the above-mentioned computer programs, computer instructions, etc. can be partitioned and stored in one or more memories. In addition, the above-mentioned computer programs, computer instructions, data, etc. can be called by the processor.

[0127] The aforementioned computer programs, computer instructions, etc. may be partitioned and stored in one or more memories, and the aforementioned computer programs, computer instructions, data, etc. may be called by a processor.

[0128] The processor is configured to execute the computer program stored in the memory to implement the various steps of the method involved in the above embodiment. For details, please refer to the relevant description in the above method embodiment.

[0129] The processor and memory can be independent structures or integrated structures. When the processor and memory are independent structures, the memory and processor can be coupled via a bus.

[0130] In an embodiment of the present disclosure, a non-temporary computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of a quality evaluation method for generating three-dimensional content from text in any of the above embodiments are implemented.

[0131] Those skilled in the art will appreciate that the embodiments of the present disclosure may be provided as methods, systems, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0132] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0133] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0135] Although the preferred embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present disclosure.

[0136] Obviously, those skilled in the art may make various changes and modifications to the present disclosure without departing from the spirit and scope of the present disclosure. Thus, if these modifications and variations of the present disclosure fall within the scope of the claims of the present disclosure and their equivalents, the present disclosure is intended to include these modifications and variations.

Claims

1. A method for evaluating the quality of three-dimensional content generated by text, characterized in that: include: Obtaining projection video and text prompts for text-generated three-dimensional content; Inputting the projection video into a pre-trained shape feature extractor to determine shape features of the text generating three-dimensional content, wherein the pre-trained shape feature extractor includes a three-dimensional encoder; Inputting the projection video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content, the pre-trained texture feature extractor comprising a first two-dimensional encoder and a second two-dimensional encoder; Inputting the projected video and the text prompt word into a pre-trained text content consistency feature extractor to determine the text content consistency feature of the text-generated three-dimensional content, wherein the pre-trained text content consistency feature extractor includes an image encoder and a text encoder; performing feature fusion processing on the shape features of the text-generated three-dimensional content, the texture features of the text-generated three-dimensional content, and the text content consistency features of the text-generated three-dimensional content to determine an evaluation score of the text-generated three-dimensional content; The step of inputting the projection video into a pre-trained shape feature extractor to determine the shape features of the text to generate three-dimensional content includes: Performing time-domain downsampling processing on the projected video to determine a downsampled video; Inputting the downsampled video into the 3D encoder to determine shape features of the 3D content generated by the text; Inputting the projection video into a pre-trained texture feature extractor to determine texture features of the text to generate three-dimensional content includes: Determining a front image and a back image of the projected video according to the projected video; Inputting the front image into the first two-dimensional encoder to determine the front texture features of the text-generated three-dimensional content; Inputting the back side image into the second two-dimensional encoder to determine the back side texture features of the text to generate three-dimensional content; fusing the front texture features of the three-dimensional content generated by the text and the back texture features of the three-dimensional content generated by the text to determine the texture features of the three-dimensional content generated by the text; The step of inputting the projected video and the text prompt words into a pre-trained text content consistency feature extractor to determine the text content consistency feature of the text-generated three-dimensional content includes: Using the first frame image of the projected video as the front image of the projected video; Inputting the front image of the projected video into the image encoder to determine image features of the text-generated three-dimensional content; Inputting the text prompt word into the text encoder to determine text features of the text to generate three-dimensional content; The image features of the text-generated three-dimensional content and the text features of the text-generated three-dimensional content are fused to determine the text content consistency features of the text-generated three-dimensional content.

2. The method according to claim 1, characterized in that The evaluation scores of the text-generated three-dimensional content include a consistency quality score, a authenticity quality score, and a perception quality score.

3. The method according to claim 1, characterized in that The determining, based on the projected video, a front image and a back image of the projected video, includes: Using the first frame image of the projected video as the front image of the projected video; The time domain center frame image of the projection video is used as the back image of the projection video.

4. The method according to claim 1, wherein The method of obtaining text to generate a projection video and text prompt words of three-dimensional content includes: Rendering the text to generate a projection video of the three-dimensional content using a rendering camera that surrounds the text to generate the three-dimensional content according to a preset projection video resolution size; Obtain preset text prompt words for generating the three-dimensional content of the text.

5. A quality evaluation system for text-generated three-dimensional content, characterized in that: include: An acquisition module, used to acquire projection video and text prompt words for generating three-dimensional content from text; a shape feature extraction module, configured to input the projection video into a pre-trained shape feature extractor to determine shape features of the text to generate three-dimensional content, wherein the pre-trained shape feature extractor includes a three-dimensional encoder; A texture feature extraction module, configured to input the projection video into a pre-trained texture feature extractor to determine texture features of the text-generated three-dimensional content, wherein the pre-trained texture feature extractor includes a first two-dimensional encoder and a second two-dimensional encoder; A text content consistency feature extraction module is used to input the projected video and the text prompt word into a pre-trained text content consistency feature extractor to determine the text content consistency feature of the three-dimensional content generated by the text, wherein the pre-trained text content consistency feature extractor includes an image encoder and a text encoder; a quality evaluation module, configured to perform feature fusion processing on shape features of the text-generated three-dimensional content, texture features of the text-generated three-dimensional content, and text content consistency features of the text-generated three-dimensional content to determine an evaluation score for the text-generated three-dimensional content; Wherein, the shape feature extraction module is used to: Performing time-domain downsampling processing on the projected video to determine a downsampled video; Inputting the downsampled video into the 3D encoder to determine shape features of the 3D content generated by the text; The texture feature extraction module is used to: Determining a front image and a back image of the projected video according to the projected video; Inputting the front image into the first two-dimensional encoder to determine the front texture features of the text-generated three-dimensional content; Inputting the back side image into the second two-dimensional encoder to determine the back side texture features of the text to generate three-dimensional content; fusing the front texture features of the three-dimensional content generated by the text and the back texture features of the three-dimensional content generated by the text to determine the texture features of the three-dimensional content generated by the text; The text content consistency feature extraction module is used to: Using the first frame image of the projected video as the front image of the projected video; Inputting the front image of the projected video into the image encoder to determine image features of the text-generated three-dimensional content; Inputting the text prompt word into the text encoder to determine text features of the text to generate three-dimensional content; The image features of the text-generated three-dimensional content and the text features of the text-generated three-dimensional content are fused to determine the text content consistency features of the text-generated three-dimensional content.

6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

7. A terminal, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Realistic three-dimensional color texture reconstruction method

    CN110599578A

  • Three-dimensional modeling method and device and server

    CN111524232A