A method, apparatus, and terminal device for generating mixed reality content based on a large AI model

By collecting and processing user voice and environmental image information, highly consistent mixed reality content is generated, which solves the problem of stiff interaction between virtual objects and reality, and improves the immersion and interactivity of the user experience.

CN119850882BActive Publication Date: 2025-06-10HARBIN SIHE INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510337545.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-10
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In existing virtual reality technologies, the interaction between virtual objects and physical objects in the real world seems stiff and unnatural, resulting in environmental coherence and user fusion perception.

Method used

By collecting the voice information and environmental image information of the target user, performing speech recognition and image generation, segmenting out object information that does not appear in the environment, virtual construction is carried out, and mixed reality content is generated to achieve a high degree of consistency between the virtual object and the real environment.

Benefits of technology

It improves the immersion and authenticity of the user experience, reduces the sense of separation between virtual and reality, and enhances the interaction between users and the virtual environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119850882B_ABST
    Figure CN119850882B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a method, apparatus, and terminal device for generating mixed reality content based on an AI large model, belonging to the field of artificial intelligence technology. The method includes: collecting initial voice information corresponding to a target user, initial image information corresponding to a target environment where the target user is located, and first three-dimensional information; performing speech recognition on the initial voice information to obtain target text information; performing image generation based on the target text information and the initial image information to obtain target image information corresponding to the target text information; performing image segmentation on the target image information according to the initial image information to obtain initial object information that does not appear in the target image information of the initial image information; performing virtual construction on the initial object information to obtain second three-dimensional information corresponding to the initial object information; and generating mixed reality content according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, and terminal device for generating mixed reality content based on an AI large model. Background Art

[0002] With the rapid progress of virtual reality technology, people have been able to reconstruct and simulate objects and scenes in the real world through this technology, providing unprecedented possibilities for various applications such as education, entertainment, medical, engineering, and other fields. However, despite the remarkable achievements of virtual reality in creating virtual environments and experiences, the existing virtual reality generated content often faces a key problem: they are usually separated from the physical objects and environments in the real world and it is difficult to directly utilize the existing objects in reality. This limitation leads to a sense of disconnection in the user experience and restricts the comprehensive application and development of virtual reality technology. In many virtual reality applications, although users can see, touch, and even interact with virtual objects, these objects lack effective association with the real environment around the users. For example, in augmented reality applications, although virtual objects may seemingly be superimposed on the real environment, their interaction with actual physical objects appears rigid and unnatural. This sense of disconnection not only disrupts the coherence of the environment but also affects the user's perception of the integration of the virtual and real worlds. Summary of the Invention

[0003] The main objective of the embodiments of the present invention is to provide a method, device, and terminal device for generating mixed reality content based on an AI large model, aiming to solve the problem in the related technology that although virtual objects are superimposed on the real environment, their interaction with actual physical objects appears rigid and unnatural, and this sense of disconnection not only disrupts the coherence of the environment but also affects the user's perception of the integration of the virtual and real worlds.

[0004] In a first aspect, the embodiments of the present invention provide a method for generating mixed reality content based on an AI large model, including:

[0005] Collecting initial voice information corresponding to a target user, initial image information corresponding to a target environment where the target user is located, and first three-dimensional information;

[0006] Performing speech recognition on the initial voice information to obtain target text information;

[0007] Generating an image according to the target text information and the initial image information to obtain target image information corresponding to the target text information;

[0008] Performing image segmentation on the target image information according to the initial image information to obtain initial object information that does not appear in the initial image information in the target image information;

[0009] Virtual construction is performed on the initial object information to obtain second three-dimensional information corresponding to the initial object information;

[0010] Based on the first three-dimensional information and the second three-dimensional information, mixed reality content generation is performed to obtain a mixed reality generation result corresponding to the initial voice information.

[0011] In a second aspect, an embodiment of the present invention provides a mixed reality content generation device based on an AI large model, including:

[0012] A data acquisition module, configured to acquire initial voice information corresponding to a target user, initial image information corresponding to a target environment where the target user is located, and first three-dimensional information;

[0013] A speech recognition module, configured to perform speech recognition on the initial voice information to obtain target text information;

[0014] An image generation module, configured to generate an image corresponding to the target text information according to the target text information and the initial image information to obtain target image information corresponding to the target text information;

[0015] An image segmentation module, configured to perform image segmentation on the target image information according to the initial image information to obtain initial object information that does not appear in the initial image information in the target image information;

[0016] A virtual construction module, configured to perform virtual construction on the initial object information to obtain second three-dimensional information corresponding to the initial object information;

[0017] A result generation module, configured to perform mixed reality content generation according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information.

[0018] In a third aspect, an embodiment of the present invention further provides a terminal device, which includes a processor, a memory, a computer program stored on the memory and executable by the processor, and a data bus for realizing connection communication between the processor and the memory. When the computer program is executed by the processor, the steps of any one of the mixed reality content generation methods based on an AI large model provided in the specification of the present invention are implemented.

[0019] An embodiment of the present invention provides a method, device, and terminal device for generating mixed reality content based on an AI large model. The method includes: collecting initial voice information of a target user, initial image information of the surrounding environment, and first three-dimensional information, then recognizing the initial voice information to obtain target text information, thereby obtaining virtual information that the target user expects to generate. Then, according to the target text information and the initial image information, image generation is performed to obtain target image information corresponding to the target text information. Then, according to the initial image information, the target image information is segmented to obtain initial object information that does not appear in the target image information, that is, an object that does not exist in the target environment where the target user is located. Then, virtual construction is performed on the initial object information to obtain second three-dimensional information corresponding to the initial object information. Then, according to the first three-dimensional information and the second three-dimensional information, mixed reality content generation is performed to obtain a mixed reality generation result corresponding to the initial voice information. This can better understand the target environment where the target user is located, making the generated mixed reality content not only visually realistic but also highly consistent with the real environment in space, providing a more immersive and real experience for the target user, reducing the sense of disconnection and detachment from the real world, greatly improving the quality and effect of the user's interaction with the digital world, and further solving the problem in the related art that although virtual objects are superimposed on the real environment, their interaction with actual physical objects is rigid and unnatural. This sense of disconnection not only destroys the coherence of the environment but also affects the user's perception of the integration of the virtual and real worlds. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic flowchart of a method for generating mixed reality content based on an AI large model provided by an embodiment of the present invention;

[0022] Figure 2 It is a schematic block diagram of the module structure of a device for generating mixed reality content based on an AI large model provided by an embodiment of the present invention;

[0023] Figure 3 It is a schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] The flowchart shown in the accompanying drawings is only an example illustration, and does not necessarily include all the content and operations / steps, nor does it necessarily execute in the described order. For example, some operations / steps can be decomposed, combined, or partially merged, so the actual execution order may change according to the actual situation.

[0026] It should be understood that the terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0027] The embodiments of the present invention provide a method, device, and terminal device for generating mixed reality content based on an AI large model. Among them, the method for generating mixed reality content based on an AI large model can be applied to a terminal device, which can be an electronic device such as a tablet computer, a notebook computer, a desktop computer, a personal digital assistant, and a wearable device. The terminal device can be a server or a server cluster.

[0028] Next, in conjunction with the accompanying drawings, some embodiments of the present invention will be described in detail. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0029] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for generating mixed reality content based on an AI large model provided by an embodiment of the present invention.

[0030] As Figure 1 shown, the method for generating mixed reality content based on an AI large model includes steps S101 to S106.

[0031] Step S101, collect the initial voice information corresponding to the target user and the initial image information and the first three-dimensional information corresponding to the target environment where the target user is located.

[0032] Exemplarily, a high-quality microphone or microphone array is used to ensure that the initial voice information of the target user can be clearly captured, and a high-resolution camera or camcorder is used to ensure that the initial image information corresponding to the target environment where the target user is located can be captured. The camera or camcorder should be placed at a position that can comprehensively cover the target environment where the target user is located.

[0033] Exemplarily, a depth sensor (such as a lidar, structured light sensor, or ToF sensor) or a stereo camera is used to ensure that the first three-dimensional information of the target environment can be captured. In addition, ensure that the acquisition times of the initial voice information, the initial image information, and the first three-dimensional information are synchronized to avoid temporal disorder among the data. A synchronization trigger or timestamp technology can be used to achieve the time alignment of multi-sensor data.

[0034] Step S102: Perform speech recognition on the initial voice information to obtain target text information.

[0035] Exemplarily, a pre-trained speech recognition model (such as a deep learning-based model, such as RNN, LSTM, Transformer, etc.) or a traditional statistical model (such as a hidden Markov model, HMM) is used to perform speech recognition on the initial voice information to obtain target text information.

[0036] In some embodiments, the performing speech recognition on the initial voice information to obtain target text information includes: performing preprocessing operations of pre-emphasis, frame windowing, and denoising on the initial voice information to obtain target voice information; using a feature extraction model to perform speech feature extraction on the obtained target voice information to obtain the initial Mel frequency cepstral coefficients, the initial short-time average energy parameters, and the initial spectral mean parameters corresponding to the target voice information; performing normalization processing on the initial Mel frequency cepstral coefficients, the initial short-time average energy parameters, and the initial spectral mean parameters to obtain target Mel frequency cepstral coefficients, target short-time average energy parameters, and target spectral mean parameters; and using a speech recognition model to perform speech recognition in combination with the target Mel frequency cepstral coefficients, the target short-time average energy parameters, and the target spectral mean parameters to obtain the target text information.

[0037] Exemplarily, a continuous initial voice signal is segmented into short-time frames (usually 20-30 milliseconds), and a window function (such as a Hamming window) is applied to each frame to reduce spectral leakage and edge effects, so as to perform a fast Fourier transform on the initial voice signal of each frame to obtain a spectral representation, and then extract the initial Mel frequency cepstral coefficients from the spectral representation, thereby capturing the spectral characteristics of the voice signal.

[0038] Exemplarily, calculate the initial short-time average energy parameter of each frame of the initial speech signal, which can reflect the energy change of the speech signal, and calculate the initial spectral mean parameter of each frame of the initial speech signal, which can capture the average level of the spectrum.

[0039] Exemplarily, calculate the mean and standard deviation of the initial Mel frequency cepstrum parameter, the initial short-time average energy parameter, and the initial spectral mean parameter respectively, and then perform mean normalization on each parameter, that is, subtract the mean of each parameter and then divide by the standard deviation to obtain the target Mel frequency cepstrum parameter, the target short-time average energy parameter, and the target spectral mean parameter. The standardization process makes all feature parameters on the same scale, avoids the problem of different dimensions between features, and improves the training effect and recognition accuracy of the model.

[0040] Exemplarily, use a speech recognition model, such as a deep learning-based model (such as RNN, LSTM, Transformer, etc.) or a traditional statistical model (such as Hidden Markov Model, HMM), and input the standardized target Mel frequency cepstrum parameter, the target short-time average energy parameter, and the target spectral mean parameter into the speech recognition model. Use the model to decode the input feature parameters to generate the corresponding target text information. The decoding process can be combined with a language model to improve the accuracy and fluency of text generation.

[0041] Specifically, combining multiple feature parameters and an efficient model can significantly improve the accuracy of speech recognition, and thus achieve an efficient and accurate conversion from the target speech information to the target text information, providing a reliable basis for subsequent image generation and mixed reality content generation.

[0042] In some embodiments, obtaining the feature extraction model includes: obtaining the training speech signal and the target feature information corresponding to the training speech signal, where the training speech signal is the audio pronunciation corresponding to each character; determining the initial model parameters corresponding to the speech recognition model, and performing operations through the input layer, hidden layer, and output layer corresponding to the initial model parameters according to the initial speech parameters and the training speech signal to obtain the predicted feature information; calculating the difference information according to the target feature information and the predicted feature information, and comparing the difference information with a preset threshold to obtain a comparison result; when the comparison result is that the difference information is greater than or equal to the preset threshold, update the initial model parameters and then continue to use the training speech signal and the target feature information to train the speech recognition model until the target model parameters corresponding to the speech recognition model are obtained; when the comparison result is that the difference information is less than the preset threshold, obtain the target model parameters corresponding to the speech recognition model, and then determine the speech recognition model according to the target model parameters.

[0043] Exemplarily, training data for training a speech recognition model is prepared, including training speech signals and their corresponding target feature information. The training speech signals are the audio pronunciations corresponding to each character. For example, the target feature information includes Mel Frequency Cepstral Coefficients (MFCC), short-time average energy, spectral mean, etc. corresponding to the training speech signals.

[0044] Exemplarily, initial model parameters corresponding to the speech recognition model are randomly generated, and then the training speech signals are input into the input layer of the model. Through the processing of the hidden layer, predicted feature information is finally obtained at the output layer, and then the difference between the target feature information and the predicted feature information is calculated, such as Mean Square Error (MSE), cross-entropy, etc. Thus, the calculated error value information is compared with a preset threshold to determine whether the error meets the requirements.

[0045] Exemplarily, when the comparison result is that the difference information is greater than or equal to the preset threshold, the initial model parameters are updated, and then the training speech signals and the target feature information are continued to be used to train the speech recognition model until the target model parameters corresponding to the speech recognition model are obtained. For example, according to the error calculation result, the initial model parameters are adjusted through the backpropagation algorithm to reduce the error information, and the forward propagation and backpropagation processes are repeated until the error information is lower than the preset threshold or a predetermined number of training times is reached.

[0046] Exemplarily, when the comparison result is that the difference information is less than the preset threshold, the target model parameters corresponding to the speech recognition model are obtained, and then the speech recognition model is determined according to the target model parameters. That is, the trained target model parameters are saved for subsequent use.

[0047] Specifically, a high-performance speech recognition model is trained, which can accurately convert speech signals into text information, providing accurate text input for subsequent image generation and mixed reality content generation.

[0048] Step S103: Image generation is performed according to the target text information and the initial image information to obtain target image information corresponding to the target text information.

[0049] Exemplarily, applying natural language processing technology to the analysis of target text information, including performing part-of-speech tagging, named entity recognition, and syntactic analysis to extract key entities in the target text information and the association relationships between the key entities. Subsequently, using models such as convolutional neural networks (CNNs) to extract features from the initial image information, such as color, texture, shape, and spatial layout, and then identifying the position information corresponding to the key entities in the initial image information. On this basis, further exploring key entities not present in the initial image information and generating relevant image information for the key entities not present in the initial image information through a target generation model. Finally, according to the identified position information and entity relationships, fusing the newly generated relevant image information with the initial image to obtain target image information that is highly consistent with the target text information. The target image information not only integrates the key entities in the target environment but also includes new key entities appearing in the target text information, thus providing support for reducing the sense of fragmentation and disconnection between virtual objects and the real world in the subsequent process.

[0050] In some embodiments, the obtaining of the target image information corresponding to the target text information by performing image generation based on the target text information and the initial image information includes: performing text-image matching on the target text information and the initial image information to obtain associated image information corresponding to the matching between the target text information and the initial image information; determining basic image information corresponding to the target text information according to the associated image information; using the text feature extraction layer of the text-to-image generation model to extract feature information from the target text information to obtain text feature information; using the image feature extraction layer of the text-to-image generation model to extract feature information from the basic image information to obtain initial image feature information; using the feature fusion layer of the text-to-image generation model to perform sentence feature fusion on the text feature information to obtain sentence feature information; using the feature addition layer of the text-to-image generation model to add random noise to the sentence feature information to obtain spliced feature information; using the text-image affine layer of the text-to-image generation model to perform image processing on the spliced feature information and the initial image feature information to obtain target image feature information; using the convolutional network layer of the text-to-image generation model to obtain intermediate image information corresponding to the target text information according to the target image feature information; and using the image enhancement layer of the text-to-image generation model to refine the intermediate image information to obtain the target image information corresponding to the target text information.

[0051] Exemplarily, using cross-modal matching technology, analyze the key elements and features in the target text information and the initial image information, determine the corresponding relationship between them, and thus extract the associated image information that matches in the target text information and the initial image information. For example, if there is a key entity "house" in the target text information, when the target text information and the initial image information match successfully, the associated image information corresponding to the house can be obtained from the initial image information.

[0052] Exemplarily, splice according to the position of the associated image information in the initial image information to obtain the basic image information corresponding to the target text information, which serves as the basis for image generation. That is, the basic image information only contains the content in which the key entity corresponding to the target text information exists in the initial image information.

[0053] Exemplarily, use the text feature extraction layer of the text-to-image generation model to perform word embedding, sentence encoding, etc. on the target text information to extract text feature information. Then, use the image feature extraction layer of the text-to-image generation model to perform operations such as convolution and pooling on the basic image information to extract the corresponding image feature information in terms of color, texture, shape, and spatial layout. Thus, use the feature fusion layer of the text-to-image generation model to fuse the extracted text feature information to form complete sentence feature information.

[0054] Exemplarily, use the feature addition layer of the text-to-image generation model to add random noise to the sentence feature information to obtain spliced feature information. Thus, through the introduction of noise, the generated image is richer in details and features, improving the visual quality.

[0055] Exemplarily, use the text-image affine layer of the text-to-image generation model to perform affine transformation and fusion processing on the spliced feature information and the initial image feature information to generate target image feature information. Furthermore, the fused image feature information is more consistent semantically and visually, improving the accuracy and quality of the generated image.

[0056] Exemplarily, use the convolutional network layer of the text-to-image generation model to generate intermediate image information through convolution operations according to the target image feature information. Through the convolution operation, the target image feature information is further strengthened, generating intermediate image information that is more in line with the target text information.

[0057] Exemplarily, use the image enhancement layer of the text-to-image generation model to perform image refinement processing on the intermediate image information to improve the image quality and details. Furthermore, through the image refinement processing, the details of the generated image are optimized, making it more in line with the description of the target text information. Thus, the target image information corresponding to the target text information is obtained.

[0058] In some embodiments, performing image processing on the spliced feature information and the initial image feature information by using the text-image affine layer of the text-to-image generation model to obtain target image feature information includes: using the relationship recognition network of the text-image affine layer to determine the target relationship corresponding to the target objects involved in the target text information based on the spliced feature information; using the position determination network of the text-image affine layer to determine the target position information corresponding to the target objects that do not exist in the initial image information according to the initial position information of the target objects in the initial image information and the target relationship; using the position determination image space layout network of the text-image affine layer to perform generation position adjustment on the initial image feature information by using the target position information to obtain the target image feature information.

[0059] Exemplarily, use the relationship recognition network in the text-image affine layer to determine the target relationship between the target objects involved in the target text information based on the spliced feature information. That is, through the relationship recognition network, the semantic relationship between the target objects involved in the target text information can be accurately captured, providing important guiding information for image generation.

[0060] Exemplarily, use the position determination network in the text-image affine layer to determine the target position information corresponding to the target objects that do not exist in the initial image information according to the initial position information of the target objects in the initial image information and the recognized target relationship. Furthermore, through the position determination network, it is ensured that the positions of the newly introduced target objects in the target image information are reasonable to conform to the relationships described in the target text information, thereby providing support for maintaining the spatial consistency between the target objects in the target image information subsequently, and thus improving the naturalness of the generated image.

[0061] Exemplarily, use the position determination image space layout network in the text-image affine layer to perform generation position adjustment on the initial image feature information by using the target position information to obtain the target image feature information. Thus, by adjusting the spatial layout of the image feature information, the generated image feature information is made more reasonable and coordinated in space, and it is ensured that the newly added target objects are coordinated with the objects in the initial image information in space, enhancing the coherence of the overall visual effect.

[0062] Specifically, through the above steps, effective fusion of the target text information and the initial image information can be achieved to generate target image feature information that not only conforms to the text description but also is consistent with the style of the initial image, thereby further generating high-quality target image information.

[0063] In some embodiments, obtaining the target image information corresponding to the target text information by refining the intermediate image information using the image enhancement layer of the text-to-image generation model includes: extracting initial image features from the intermediate image information using the image feature recognition network of the image enhancement layer; calculating an information correlation value between the spliced feature information and the initial image features using the text-image correlation network of the image enhancement layer; determining the target weight corresponding to the initial image features according to the information correlation value using the weight calculation layer of the image enhancement layer; updating the initial image features according to the target weight using the feature update network of the image enhancement layer to obtain updated target image features; supplementing missing features of the target image features using the detail correction network of the image enhancement layer to obtain final image features; generating the target image information corresponding to the target text information using the image generation network of the image enhancement layer; wherein, the information correlation value is obtained according to the following formula:

[0064]

[0065] Wherein, represents the information correlation value between the i-th spliced feature information and the initial image features, represents the i-th spliced feature information, represents the matrix required to transform the i-th spliced feature information, represents the t-th initial image feature, and m represents the number of features corresponding to the initial image features, represents the transformation required matrix, represents an activation function, and represent 1*1 convolutions for dimensionality conversion so that the i-th spliced feature information and have the same feature space.

[0066] Exemplarily, the image feature recognition network in the image enhancement layer is used to extract image features from the intermediate image information to obtain initial image features. Then, the text-image correlation network in the image enhancement layer is used to calculate the information correlation value between the spliced feature information and the initial image features. By calculating the information correlation value, the correlation between the text and the image features is quantified, providing a basis for weight calculation and helping to determine which image features need to be strengthened or weakened to better match the text description. Wherein, the information correlation value is obtained according to the following formula:

[0067]

[0068] Wherein, Represents the information association value between the i-th splicing feature information and the initial image feature, Represents the i-th splicing feature information, Represents the matrix required to transform the i-th splicing feature information, Represents the t-th initial image feature, and m represents the number of features corresponding to the initial image feature, Represents the transformation The required matrix, Represents the activation function, And Represents a 1*1 convolution for dimensionality conversion so that the i-th splicing feature information And Have the same feature space.

[0069] Exemplarily, by calculating the specific information association value, the correlation degree between the splicing feature information and the initial image feature can be clearly understood, thereby providing a specific reference basis for subsequent weight calculation and making the feature update more accurate. The non-linear characteristic of the activation function can capture the complex non-linear relationship between the splicing feature information and the initial image feature, improve the accuracy of the information association value, and then enhance the expression ability of the model through non-linear transformation, enabling it to better process complex data.

[0070] Exemplarily, the weight calculation layer in the image enhancement layer is used to perform data normalization according to the information association value to calculate the target weight corresponding to the initial image feature, and then the feature update network in the image enhancement layer is used to perform feature update on the initial image feature based on the target weight to obtain the updated target image feature.

[0071] Exemplarily, the detail correction network in the image enhancement layer is used to supplement the missing features of the target image feature to obtain the final image feature. The detail correction network supplements the missing features, making the generated image more complete and rich, thereby making the image more real and delicate. Finally, the image generation network in the image enhancement layer is used to generate the target image information based on the final image feature, and then the generated image precisely matches the target text information, realizing high-quality text-to-image conversion. Thus, high-quality generation from the intermediate image information to the final target image information can be achieved, ensuring that the generated image not only conforms to the text description but also has a high degree of realism and delicacy visually.

[0072] Step S104, perform image segmentation on the target image information according to the initial image information to obtain the initial object information that does not appear in the initial image information of the target image information.

[0073] Exemplarily, target segmentation is performed on the initial image information to obtain a plurality of first segmentation objects, and type recognition is performed on each first segmentation object to obtain a first target type. Target segmentation is performed on the target image information to obtain a plurality of second segmentation objects, and type recognition is performed on each second segmentation object to obtain a second target type.

[0074] Exemplarily, the first target type and the second target type are subjected to an intersection process to determine the relationship between the two. Specifically, by performing an intersection operation on the first target type and the second target type, it can be identified which second segmentation objects in the second target type already exist in the first target type and which second segmentation objects do not exist in the first target type. When some second segmentation objects in the second target type do not exist in the first target type, the initial object information of these second segmentation objects, that is, the initial image information, will be extracted as the key input for the subsequent generation stage. Since there is no corresponding first three-dimensional information for these initial object information in the first target type, when generating the mixed reality generation result, it is necessary to specifically generate corresponding second three-dimensional information for these initial object information. In this way, by generating these missing second three-dimensional information, it can be ensured that the finally generated mixed reality result can be perfectly integrated with the real environment, providing a more realistic and coherent user experience. This process not only fills the data gap but also provides a solid foundation for generating mixed reality content highly consistent with the real environment.

[0075] Step S105: Virtually construct the initial object information to obtain the second three-dimensional information corresponding to the initial object information.

[0076] Exemplarily, using a feature extraction algorithm or model, key features such as shape, texture, and material are extracted from the initial object information to provide basic data for virtual construction. Then, according to the characteristics of the initial object information, a suitable three-dimensional modeling technology is selected, such as a generative model based on deep learning, a traditional three-dimensional reconstruction algorithm, etc. Then, the parameters of the model are adjusted according to the requirements, such as resolution, level of detail, lighting conditions, etc., to obtain an ideal three-dimensional effect. Then, the extracted features are input into the selected model, and the generation process is executed to output the second three-dimensional information corresponding to the initial object information.

[0077] In some embodiments, virtual construction of the initial object information to obtain the second three-dimensional information corresponding to the initial object information includes: performing filtering processing on the initial object information to obtain a target filtering result, and performing color space conversion on the initial object information to obtain target color information corresponding to the target color space; determining a target segmentation block, and performing image segmentation on the initial object information according to the target segmentation block to obtain a target segmentation result corresponding to the target segmentation block; determining a target feature vector corresponding to the target segmentation result according to the target filtering result and the target color information; using the brightness enhancement layer of the image reconstruction model to perform feature brightness enhancement processing on the target feature vector to obtain a first enhanced feature vector; using the detail reconstruction layer of the image reconstruction model to perform detail feature enhancement processing on the first enhanced feature vector to obtain a second enhanced feature vector; using the target generation layer of the image reconstruction model to perform image generation on the second enhanced feature vector to obtain target object information; and performing virtual construction according to the target object information to obtain the second three-dimensional information.

[0078] Exemplarily, a filtering algorithm (such as Gaussian filtering, median filtering, etc.) is used to process the initial object information to obtain a target filtering result, thereby removing unnecessary details and reducing interference information in subsequent processing.

[0079] Exemplarily, a color space conversion algorithm (such as RGB to HSV, Lab, etc.) is used for conversion to obtain target color information, thereby converting to a more suitable color space, which can better extract and analyze color information, and thus provide richer color features, contributing to subsequent feature extraction and image segmentation.

[0080] Exemplarily, the initial object information is segmented according to the target segmentation block to obtain a target segmentation result under the target segmentation block. The size of the target segmentation block can be 2*2 or 4*4.

[0081] Exemplarily, a feature extraction algorithm (such as HOG, SIFT, CNN, etc.) is used to determine a target feature vector corresponding to the target segmentation result according to the target filtering result and the target color information. Then, the brightness enhancement layer of the image reconstruction model is used to perform feature brightness enhancement processing on the target feature vector to obtain a first enhanced feature vector. For example, a brightness enhancement algorithm (such as histogram equalization, contrast adjustment, etc.) is used for processing, thereby improving the visual quality of the image and contributing to clearer image generation.

[0082] Exemplarily, using the detail reconstruction layer of the image reconstruction model, detail feature enhancement processing is performed according to the first enhanced feature vector to obtain a second enhanced feature vector. For example, processing is carried out using detail enhancement algorithms (such as sharpening, detail filtering, etc.), thereby enhancing the detail information of the image, making the generated image more real and delicate, and improving the detail quality of the image, making the image more in line with expectations.

[0083] Exemplarily, using the target generation layer of the image reconstruction model, image generation is performed on the second enhanced feature vector to obtain target object information. For example, generation is carried out using image generation algorithms (such as GAN, VAE, etc.), and then high-quality images are generated based on the enhanced feature vector to ensure the realism and details of the image.

[0084] Exemplarily, virtual construction of the target object information is performed using 3D modeling techniques (such as Mesh generation, point cloud reconstruction, etc.) to obtain corresponding second 3D information, so that the generated second 3D information can truly reflect the target object.

[0085] Specifically, by effectively converting the initial object information into high-quality second 3D information, it is ensured that the generated 3D model not only meets the requirements of the initial object information but also has a high degree of realism and detail quality. This provides a solid data foundation for the subsequent mixed reality generation results, ensuring that the generated content can be perfectly integrated with the real environment and enhancing the user experience.

[0086] Step S106: Generate mixed reality content according to the first 3D information and the second 3D information to obtain the mixed reality generation result corresponding to the initial voice information.

[0087] Exemplarily, techniques such as coordinate system conversion and scale normalization are used to align the first 3D information and the second 3D information, and then a feature fusion algorithm, such as weighted fusion, deep learning fusion, etc., is used to integrate the features of the aligned first 3D information and the second 3D information to obtain target 3D information. Furthermore, a 3D rendering engine, such as OpenGL, Unity, is used to render the fused target 3D information according to the color of the first 3D information in the initial image information and the color of the second 3D image in the target image information to generate a visual mixed reality generation result.

[0088] Through these steps, the first 3D information and the second 3D information can be effectively fused, combined with the initial voice information, to generate high-quality mixed reality content, providing users with a real and immersive experience. This not only improves the quality of the mixed reality content but also enhances the interactivity between users and the virtual environment, improving the user experience.

[0089] In some embodiments, generating the mixed reality generation result corresponding to the initial voice information according to the first three-dimensional information and the second three-dimensional information includes: splicing the first three-dimensional information and the second three-dimensional information to obtain the initial coincidence information corresponding to the first three-dimensional information and the second three-dimensional information; performing information fusion on the initial coincidence information to obtain the target coincidence information corresponding to the initial coincidence information; and generating mixed reality content according to the first three-dimensional information, the second three-dimensional information, and the target coincidence information to obtain the mixed reality generation result corresponding to the initial voice information.

[0090] Exemplarily, the first three-dimensional information and the second three-dimensional information are spliced to obtain initial coincidence information. For example, through coordinate alignment and data splicing techniques, the first three-dimensional information and the second three-dimensional information are spliced in space to form a unified three-dimensional data set.

[0091] Exemplarily, an information fusion algorithm (such as weighted average, deep learning fusion, etc.) is used to process the spliced initial coincidence information to eliminate redundant information and enhance key features, thereby obtaining the target coincidence information.

[0092] Exemplarily, the initial coincidence information is removed from the first three-dimensional information to obtain the third three-dimensional information, and the initial coincidence information is removed from the second three-dimensional information to obtain the fourth three-dimensional information. Then, data merging is performed according to the third three-dimensional information, the fourth three-dimensional information, and the target coincidence information to obtain the target three-dimensional information.

[0093] Exemplarily, a three-dimensional rendering engine, such as OpenGL or Unity, is used for the fused target three-dimensional information to perform rendering according to the color of the third three-dimensional information in the initial image information, the color of the fourth three-dimensional image in the target image information, and the color after average color processing of the target coincidence information in the initial image information and the target image information, generating a visual mixed reality generation result. Thus, by fusing multi-source three-dimensional information, the generated mixed reality content is more real and delicate and can be seamlessly integrated with the real environment.

[0094] Specifically, through information splicing, fusion, and mixed reality content generation, the first three-dimensional information and the second three-dimensional information can be integrated into high-quality mixed reality content. This process not only improves the integrity and accuracy of the information but also enhances the realism and immersion of the mixed reality content, providing users with a richer and more real experience. The final mixed reality generation result can be perfectly integrated with the real environment to meet the requirements of various application scenarios.

[0095] Please refer to Figure 2 , Figure 2A hybrid reality content generation device 200 based on an AI large model provided by an embodiment of the present application. The hybrid reality content generation device 200 based on the AI large model includes a data acquisition module 201, a speech recognition module 202, an image generation module 203, an image segmentation module 204, a virtual construction module 205, and a result generation module 206. Among them, the data acquisition module 201 is configured to collect initial speech information corresponding to a target user, initial image information corresponding to a target environment where the target user is located, and first three-dimensional information; the speech recognition module 202 is configured to perform speech recognition on the initial speech information to obtain target text information; the image generation module 203 is configured to generate an image according to the target text information and the initial image information to obtain target image information corresponding to the target text information; the image segmentation module 204 is configured to perform image segmentation on the target image information according to the initial image information to obtain initial object information that does not appear in the initial image information in the target image information; the virtual construction module 205 is configured to perform virtual construction on the initial object information to obtain second three-dimensional information corresponding to the initial object information; the result generation module 206 is configured to generate hybrid reality content according to the first three-dimensional information and the second three-dimensional information to obtain a hybrid reality generation result corresponding to the initial speech information.

[0096] In some embodiments, during the process of performing speech recognition on the initial speech information to obtain target text information, the speech recognition module 202 performs:

[0097] Performing preprocessing operations of pre-emphasis, frame addition with windowing, and denoising on the initial speech information to obtain target speech information;

[0098] Using a feature extraction model to perform speech feature extraction on the obtained target speech information to obtain initial Mel frequency cepstrum parameters, initial short-time average energy parameters, and initial spectral mean parameters corresponding to the target speech information;

[0099] Performing normalization processing on the initial Mel frequency cepstrum parameters, the initial short-time average energy parameters, and the initial spectral mean parameters to obtain target Mel frequency cepstrum parameters, target short-time average energy parameters, and target spectral mean parameters;

[0100] Using a speech recognition model to perform speech recognition in combination with the target Mel frequency cepstrum parameters, the target short-time average energy parameters, and the target spectral mean parameters to obtain the target text information.

[0101] In some embodiments, during the process of obtaining the feature extraction model, the speech recognition module 202 performs:

[0102] Obtain a training speech signal and target feature information corresponding to the training speech signal, where the training speech signal is the audio pronunciation corresponding to each character;

[0103] Determine initial model parameters corresponding to the speech recognition model, and perform operations through the input layer, hidden layer, and output layer corresponding to the initial model parameters based on the initial speech parameters and the training speech signal to obtain predicted feature information;

[0104] Perform difference calculation based on the target feature information and the predicted feature information to obtain difference information, and compare the difference information with a preset threshold to obtain a comparison result;

[0105] When the comparison result is that the difference information is greater than or equal to the preset threshold, update the initial model parameters and then continue to perform model training on the speech recognition model using the training speech signal and the target feature information until the target model parameters corresponding to the speech recognition model are obtained;

[0106] When the comparison result is that the difference information is less than the preset threshold, obtain the target model parameters corresponding to the speech recognition model, and then determine the speech recognition model based on the target model parameters.

[0107] In some embodiments, during the process of generating a target image information corresponding to the target text information by the image generation module 203 according to the target text information and the initial image information, the following is performed:

[0108] Perform text-image matching on the target text information and the initial image information to obtain associated image information corresponding to the matching of the target text information and the initial image information;

[0109] Determine basic image information corresponding to the target text information according to the associated image information;

[0110] Use the text feature extraction layer of the text-to-image generation model to extract feature information from the target text information to obtain text feature information;

[0111] Use the image feature extraction layer of the text-to-image generation model to extract feature information from the basic image information to obtain initial image feature information;

[0112] Use the feature fusion layer of the text-to-image generation model to perform sentence feature fusion on the text feature information to obtain sentence feature information;

[0113] Use the feature addition layer of the text-to-image generation model to add random noise to the sentence feature information to obtain spliced feature information;

[0114] Use the text-image affine layer of the text-to-image generation model to perform image processing on the spliced feature information and the initial image feature information to obtain target image feature information;

[0115] Use the convolutional network layer of the text-to-image generation model to obtain intermediate image information corresponding to the target text information according to the target image feature information;

[0116] Use the image enhancement layer of the text-to-image generation model to perform image refinement on the intermediate image information to obtain the target image information corresponding to the target text information.

[0117] In some embodiments, during the process of using the text-image affine layer of the text-to-image generation model to perform image processing on the spliced feature information and the initial image feature information to obtain target image feature information, the image generation module 203 performs:

[0118] Use the relationship recognition network of the text-image affine layer to determine the target relationship corresponding to the target objects involved in the target text information based on the spliced feature information;

[0119] Use the position determination network of the text-image affine layer to determine the target position information corresponding to the target objects that do not exist in the initial image information according to the initial position information of the target objects in the initial image information and the target relationship;

[0120] Use the position determination image space layout network of the text-image affine layer to perform generation position adjustment on the initial image feature information by using the target position information to obtain the target image feature information.

[0121] In some embodiments, during the process of using the image enhancement layer of the text-to-image generation model to perform image refinement on the intermediate image information to obtain the target image information corresponding to the target text information, the image generation module 203 performs:

[0122] Use the image feature recognition network of the image enhancement layer to perform image feature extraction on the intermediate image information to obtain initial image features;

[0123] Use the text-image association network of the image enhancement layer to calculate the information association value between the spliced feature information and the initial image features;

[0124] Use the weight calculation layer of the image enhancement layer to determine the target weight corresponding to the initial image features according to the information association value;

[0125] Use the feature update network of the image enhancement layer to perform feature update on the initial image features according to the target weight to obtain the updated target image features;

[0126] Use the detail correction network of the image enhancement layer to supplement the missing features of the target image features to obtain the final image features;

[0127] Use the image generation network of the image enhancement layer to generate the target image information corresponding to the target text information according to the final image features;

[0128] Among them, the information association value is obtained according to the following formula:

[0129]

[0130] Among them, represents the information association value between the i-th splicing feature information and the initial image features, represents the i-th splicing feature information, represents the matrix required to transform the i-th splicing feature information, represents the t-th initial image feature, and m represents the number of features corresponding to the initial image features, represents the transformation required matrix, represents the activation function, and represent 1*1 convolution, which is used for dimensionality conversion so that the i-th splicing feature information and have the same feature space.

[0131] In some embodiments, during the process of virtual construction of the initial object information to obtain the second three-dimensional information corresponding to the initial object information, the virtual construction module 205 performs:

[0132] Perform filtering processing on the initial object information to obtain a target filtering result, and perform color space conversion on the initial object information to obtain target color information corresponding to the target color space;

[0133] Determine the target segmentation block, and perform image segmentation on the initial object information according to the target segmentation block to obtain the target segmentation result corresponding to the target segmentation block;

[0134] Determine the target feature vector corresponding to the target segmentation result according to the target filtering result and the target color information;

[0135] Use the brightness enhancement layer of the image reconstruction model to perform feature brightness enhancement processing on the target feature vector to obtain the first enhanced feature vector;

[0136] The detail reconstruction layer of the image reconstruction model performs detail feature enhancement processing on the basis of the first enhanced feature vector to obtain a second enhanced feature vector;

[0137] The target generation layer of the image reconstruction model performs image generation on the second enhanced feature vector to obtain target object information;

[0138] Virtual construction is performed according to the target object information to obtain the second three-dimensional information.

[0139] In some embodiments, during the process of generating the mixed reality content corresponding to the initial voice information according to the first three-dimensional information and the second three-dimensional information, the result generation module 206 performs the following:

[0140] The first three-dimensional information and the second three-dimensional information are spliced to obtain initial coincidence information corresponding to the first three-dimensional information and the second three-dimensional information;

[0141] The initial coincidence information is fused to obtain target coincidence information corresponding to the initial coincidence information;

[0142] Mixed reality content is generated according to the first three-dimensional information, the second three-dimensional information, and the target coincidence information to obtain the mixed reality generation result corresponding to the initial voice information.

[0143] In some embodiments, the mixed reality content generation device 200 based on the AI large model can be applied to a terminal device.

[0144] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described mixed reality content generation device 200 based on the AI large model can refer to the corresponding process in the foregoing embodiment of the mixed reality content generation method based on the AI large model, and will not be elaborated herein.

[0145] Please refer to Figure 3 , Figure 3 which is a schematic block diagram of the structure of a terminal device provided by an embodiment of the present invention.

[0146] As Figure 3 shown, the terminal device 300 includes a processor 301 and a memory 302, and the processor 301 and the memory 302 are connected through a bus 303, and this bus is, for example, an I2C (Inter-integrated Circuit) bus.

[0147] Specifically, the processor 301 is used to provide computing and control capabilities to support the operation of the entire terminal device. The processor 301 can be a central processing unit (CPU), and this processor 301 can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or this processor can also be any conventional processor, etc.

[0148] Specifically, the memory 302 can be a Flash chip, read-only memory (ROM), magnetic disk, optical disc, USB flash drive, or mobile hard disk, etc.

[0149] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of the embodiment of the present invention, and does not constitute a limitation on the terminal device to which the solution of the embodiment of the present invention is applied. A specific server may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0150] Among them, the processor is used to run the computer program stored in the memory and implement any one of the AI large model-based hybrid reality content generation methods provided by the embodiments of the present invention when executing the computer program.

[0151] In one embodiment, the processor is used to run the computer program stored in the memory and implement the following steps when executing the computer program:

[0152] Collect the initial voice information corresponding to the target user, the initial image information and the first three-dimensional information corresponding to the target environment where the target user is located;

[0153] Perform speech recognition on the initial voice information to obtain target text information;

[0154] Generate an image according to the target text information and the initial image information to obtain target image information corresponding to the target text information;

[0155] Perform image segmentation on the target image information according to the initial image information to obtain initial object information that does not appear in the initial image information in the target image information;

[0156] Perform virtual construction on the initial object information to obtain the second three-dimensional information corresponding to the initial object information;

[0157] Generate mixed reality content based on the first three-dimensional information and the second three-dimensional information to obtain the mixed reality generation result corresponding to the initial voice information.

[0158] In some embodiments, during the process of performing speech recognition on the initial voice information to obtain the target text information, the processor 301 executes:

[0159] Perform preprocessing operations of pre-emphasis, frame windowing, and denoising on the initial voice information to obtain the target voice information;

[0160] Use a feature extraction model to perform speech feature extraction on the obtained target voice information to obtain the initial Mel frequency cepstrum parameters, initial short-time average energy parameters, and initial spectral mean parameters corresponding to the target voice information;

[0161] Perform normalization processing on the initial Mel frequency cepstrum parameters, the initial short-time average energy parameters, and the initial spectral mean parameters to obtain the target Mel frequency cepstrum parameters, target short-time average energy parameters, and target spectral mean parameters;

[0162] Use a speech recognition model to perform speech recognition in combination with the target Mel frequency cepstrum parameters, the target short-time average energy parameters, and the target spectral mean parameters to obtain the target text information.

[0163] In some embodiments, during the process of obtaining the feature extraction model, the processor 301 executes:

[0164] Obtain the training speech signal and the target feature information corresponding to the training speech signal, where the training speech signal is the audio pronunciation corresponding to each character;

[0165] Determine the initial model parameters corresponding to the speech recognition model, and perform operations through the input layer, hidden layer, and output layer corresponding to the initial model parameters according to the initial speech parameters and the training speech signal to obtain the predicted feature information;

[0166] Perform difference calculation according to the target feature information and the predicted feature information to obtain the difference information, and compare the difference information with a preset threshold to obtain a comparison result;

[0167] When the comparison result is that the difference information is greater than or equal to the preset threshold, update the initial model parameters and then continue to perform model training on the speech recognition model using the training speech signal and the target feature information until the target model parameters corresponding to the speech recognition model are obtained.

[0168] When the comparison result indicates that the difference information is less than the preset threshold, the target model parameters corresponding to the speech recognition model are obtained, and then the speech recognition model is determined according to the target model parameters.

[0169] In some embodiments, during the process of generating the target image information corresponding to the target text information by the processor 301 based on the target text information and the initial image information, the following operations are performed:

[0170] Perform text-image matching on the target text information and the initial image information to obtain the associated image information corresponding to the matching of the target text information and the initial image information;

[0171] Determine the basic image information corresponding to the target text information according to the associated image information;

[0172] Use the text feature extraction layer of the text-to-image generation model to extract feature information from the target text information to obtain text feature information;

[0173] Use the image feature extraction layer of the text-to-image generation model to extract feature information from the basic image information to obtain initial image feature information;

[0174] Use the feature fusion layer of the text-to-image generation model to perform sentence feature fusion on the text feature information to obtain sentence feature information;

[0175] Use the feature addition layer of the text-to-image generation model to add random noise to the sentence feature information to obtain spliced feature information;

[0176] Use the text-image affine layer of the text-to-image generation model to perform image processing on the spliced feature information and the initial image feature information to obtain target image feature information;

[0177] Use the convolutional network layer of the text-to-image generation model to obtain intermediate image information corresponding to the target text information according to the target image feature information;

[0178] Use the image enhancement layer of the text-to-image generation model to refine the intermediate image information to obtain the target image information corresponding to the target text information.

[0179] In some embodiments, during the process of using the text-image affine layer of the text-to-image generation model to perform image processing on the spliced feature information and the initial image feature information to obtain target image feature information, the following operations are performed:

[0180] The relationship recognition network using the text image affine layer determines the target relationship corresponding to the target objects involved in the target text information based on the splicing feature information;

[0181] The position determination network using the text image affine layer determines the target position information corresponding to the target objects that do not exist in the initial image information according to the initial position information of the target objects in the initial image information and the target relationship;

[0182] The position determination image space layout network using the text image affine layer performs a generation position adjustment on the initial image feature information using the target position information to obtain the target image feature information.

[0183] In some embodiments, when the processor 301 performs image refinement on the intermediate image information using the image enhancement layer of the text generation image model to obtain the target image information corresponding to the target text information, it executes:

[0184] The image feature recognition network using the image enhancement layer extracts image features from the intermediate image information to obtain initial image features;

[0185] The text-image association network using the image enhancement layer calculates the information association value between the splicing feature information and the initial image features;

[0186] The weight calculation layer using the image enhancement layer determines the target weight corresponding to the initial image features according to the information association value;

[0187] The feature update network using the image enhancement layer updates the features of the initial image features according to the target weight to obtain the updated target image features;

[0188] The detail correction network using the image enhancement layer supplements the missing features of the target image features to obtain the final image features;

[0189] The image generation network using the image enhancement layer generates the target image information corresponding to the target text information according to the final image features;

[0190] Among them, the information association value is obtained according to the following formula:

[0191]

[0192] Among them, represents the information association value between the i-th splicing feature information and the initial image features, represents the i-th splicing feature information, represents the matrix required for converting the i-th piecewise feature information represents the t-th initial image feature, and m represents the number of features corresponding to the initial image feature represents the conversion required matrix represents the activation function and represents a 1×1 convolution for dimensionality conversion so that the i-th piecewise feature information and have the same feature space

[0193] In some embodiments, during the process of virtually constructing the initial object information to obtain the second three-dimensional information corresponding to the initial object information, the processor 301 executes:

[0194] Performs filtering processing on the initial object information to obtain a target filtering result, and performs color space conversion on the initial object information to obtain target color information corresponding to the target color space;

[0195] Determines a target segmentation block, and performs image segmentation on the initial object information according to the target segmentation block to obtain a target segmentation result corresponding to the target segmentation block;

[0196] Determines a target feature vector corresponding to the target segmentation result according to the target filtering result and the target color information;

[0197] Uses the brightness enhancement layer of the image reconstruction model to perform feature brightness enhancement processing on the target feature vector to obtain a first enhanced feature vector;

[0198] Uses the detail reconstruction layer of the image reconstruction model to perform detail feature enhancement processing on the first enhanced feature vector to obtain a second enhanced feature vector;

[0199] Uses the target generation layer of the image reconstruction model to perform image generation on the second enhanced feature vector to obtain target object information;

[0200] Performs virtual construction according to the target object information to obtain the second three-dimensional information

[0201] In some embodiments, during the process of generating mixed reality content according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information, the processor 301 executes:

[0202] Performs information splicing on the first three-dimensional information and the second three-dimensional information to obtain initial coincidence information corresponding to the first three-dimensional information and the second three-dimensional information;

[0203] Perform information fusion on the initial coincidence information to obtain the target coincidence information corresponding to the initial coincidence information;

[0204] Generate mixed reality content based on the first 3D information, the second 3D information, and the target coincidence information to obtain the mixed reality generation result corresponding to the initial voice information.

[0205] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described terminal device can refer to the corresponding process in the foregoing embodiment of the method for generating mixed reality content based on the AI large model, and will not be elaborated herein.

[0206] The embodiment of the present invention further provides a storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any one of the methods for generating mixed reality content based on the AI large model provided in the specification of the embodiment of the present invention.

[0207] Among them, the storage medium can be an internal storage unit of the terminal device in the foregoing embodiment, such as the hard disk or memory of the terminal device. The storage medium can also be an external storage device of the terminal device, such as a plug-in hard disk equipped on the terminal device, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc.

[0208] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations. In the hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation. Some or all physical components can be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include a computer storage medium (or non-transitory medium) and a communication medium (or transitory medium). As is well known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory, or other memory technologies, CD-ROM, digital versatile disk (DVD), or other optical disk storage, magnetic cassette, tape, disk storage, or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, as is well known to those of ordinary skill in the art, a communication medium typically contains computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transmission mechanism, and can include any information delivery medium.

[0209] It should be understood that the term "and / or" used in the specification and appended claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations. It should be noted that in this text, the term "comprises," "comprising," or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article, or system that comprises a series of elements includes not only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or system. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article, or system that comprises the element.

[0210] The serial numbers of the embodiments of the present invention above are only for description and do not represent the superiority or inferiority of the embodiments. The above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for generating mixed reality content based on an AI big model, characterized in that: The method comprises: Collecting initial voice information corresponding to a target user and initial image information and first three-dimensional information corresponding to a target environment where the target user is located; Performing speech recognition on the initial speech information to obtain target text information; Performing image generation according to the target text information and the initial image information to obtain target image information corresponding to the target text information; Performing image segmentation on the target image information according to the initial image information to obtain initial object information of the target image information that does not appear in the initial image information; Performing virtual construction on the initial object information to obtain second three-dimensional information corresponding to the initial object information; Generate mixed reality content according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information; The step of performing image generation according to the target text information and the initial image information to obtain the target image information corresponding to the target text information includes: Performing text-image matching on the target text information and the initial image information to obtain associated image information corresponding to the matching of the target text information and the initial image information; Determining basic image information corresponding to the target text information according to the associated image information; Using the text feature extraction layer of the text generation image model to extract features of the target text information to obtain text feature information; Using the image feature extraction layer of the text-generated image model to extract features from the basic image information to obtain initial image feature information; Using the feature fusion layer of the text-to-image model to perform sentence feature fusion on the text feature information to obtain sentence feature information; Using the feature adding layer of the text-to-image model, random noise is added to the sentence feature information to obtain splicing feature information; Using the text image affine layer of the text generation image model to perform image processing on the splicing feature information and the initial image feature information to obtain target image feature information; Obtaining intermediate image information corresponding to the target text information according to the target image feature information using the convolutional network layer of the text-to-image model; Using the image enhancement layer of the text generation image model to perform image refinement on the intermediate image information to obtain the target image information corresponding to the target text information; The step of performing image refinement on the intermediate image information using the image enhancement layer of the text-generated image model to obtain the target image information corresponding to the target text information includes: Calculating the information association value between the splicing feature information and the initial image feature through the text-image association network of the image enhancement layer; The information association value is obtained according to the following formula: ; in, represents the information association value between the i-th splicing feature information and the initial image feature, represents the i-th splicing feature information, represents the matrix required to transform the i-th splicing feature information, represents the tth initial image feature, m represents the number of features corresponding to the initial image feature, Representation conversion The required matrix, represents the activation function, and Represents a 1*1 convolution, which is used to perform dimensional conversion so that the i-th concatenated feature information and have the same feature space.

2. The method according to claim 1, characterized in that The performing speech recognition on the initial speech information to obtain target text information includes: The target voice information is obtained after pre-processing operations of pre-emphasis, frame windowing and denoising are performed on the initial voice information; Performing speech feature extraction on the target speech information using a feature extraction model to obtain initial Mel frequency cepstrum parameters, initial short-time average energy parameters, and initial spectrum mean parameters corresponding to the target speech information; Standardizing the initial Mel frequency cepstrum parameter, the initial short-time average energy parameter and the initial spectrum mean parameter to obtain target Mel frequency cepstrum parameter, target short-time average energy parameter and target spectrum mean parameter; The target text information is obtained by performing speech recognition using a speech recognition model in combination with the target Mel-frequency cepstrum parameter, the target short-time average energy parameter and the target spectrum mean parameter.

3. The method according to claim 2, characterized in that Obtaining the feature extraction model includes: Obtaining a training speech signal and target feature information corresponding to the training speech signal, wherein the training speech signal corresponds to an audio pronunciation of each character; Determine initial model parameters corresponding to the speech recognition model, and perform operations through an input layer, a hidden layer, and an output layer corresponding to the initial model parameters according to the initial model parameters and the training speech signal to obtain prediction feature information; Performing a difference calculation based on the target feature information and the predicted feature information to obtain difference information, and comparing the difference information with a preset threshold to obtain a comparison result; When the comparison result is that the difference information is greater than or equal to the preset threshold, the initial model parameters are updated and then the speech recognition model is continuously trained using the training speech signal and the target feature information until the target model parameters corresponding to the speech recognition model are obtained; When the comparison result is that the difference information is less than the preset threshold, the target model parameters corresponding to the speech recognition model are obtained, and then the speech recognition model is determined according to the target model parameters.

4. The method according to claim 1, characterized in that: The step of performing image processing on the splicing feature information and the initial image feature information using the text image affine layer of the text generation image model to obtain target image feature information includes: Determine the corresponding target relationship between the target objects involved in the target text information based on the splicing feature information using the relationship recognition network of the text image affine layer; Determine the target position information corresponding to the target object that does not exist in the initial image information according to the initial position information of the target object in the initial image information and the target relationship using the position determination network of the text image affine layer; The image space layout network is determined by utilizing the position of the text image affine layer, and the target position information is used to adjust the generation position of the initial image feature information to obtain the target image feature information.

5. The method according to claim 1, characterized in that The step of performing image refinement on the intermediate image information by using the image enhancement layer of the text generation image model to obtain the target image information corresponding to the target text information includes: Using the image feature recognition network of the image enhancement layer to extract image features from the intermediate image information to obtain the initial image features; Determine the target weight corresponding to the initial image feature according to the information association value using the weight calculation layer of the image enhancement layer; Using the feature update network of the image enhancement layer to update the initial image features according to the target weights to obtain updated target image features; Using the detail correction network of the image enhancement layer to supplement the missing features of the target image features to obtain the final image features; The image generation network of the image enhancement layer is used to generate the target image information corresponding to the target text information according to the final image features.

6. The method according to claim 1, characterized in that The virtually constructing the initial object information to obtain the second three-dimensional information corresponding to the initial object information includes: Performing filtering processing on the initial object information to obtain a target filtering result, and performing color space conversion on the initial object information to obtain corresponding target color information in a target color space; Determine a target segmentation block, and perform image segmentation on the initial object information according to the target segmentation block to obtain a target segmentation result corresponding to the target segmentation block; Determine a target feature vector corresponding to the target segmentation result according to the target filtering result and the target color information; Using the brightness enhancement layer of the image reconstruction model to perform feature brightness enhancement processing according to the target feature vector to obtain a first enhanced feature vector; Using the detail reconstruction layer of the image reconstruction model, performing detail feature enhancement processing according to the first enhanced feature vector to obtain a second enhanced feature vector; Using the target generation layer of the image reconstruction model to perform image generation on the second enhanced feature vector to obtain target object information; The second three-dimensional information is obtained by performing virtual construction according to the target object information.

7. The method according to claim 1, characterized in that The generating mixed reality content according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information includes: Splicing the first three-dimensional information and the second three-dimensional information to obtain initial overlap information corresponding to the first three-dimensional information and the second three-dimensional information; Performing information fusion on the initial overlap information to obtain target overlap information corresponding to the initial overlap information; Mixed reality content is generated according to the first three-dimensional information, the second three-dimensional information and the target coincidence information to obtain the mixed reality generation result corresponding to the initial voice information.

8. A mixed reality content generation device based on an AI big model, characterized in that: include: A data acquisition module, used for acquiring initial voice information corresponding to a target user and initial image information and first three-dimensional information corresponding to a target environment where the target user is located; A speech recognition module, used for performing speech recognition on the initial speech information to obtain target text information; An image generation module is used to perform image generation based on the target text information and the initial image information to obtain target image information corresponding to the target text information; wherein, the image generation based on the target text information and the initial image information to obtain target image information corresponding to the target text information includes: performing text-image matching on the target text information and the initial image information to obtain associated image information corresponding to the target text information when the target text information matches the initial image information; determining basic image information corresponding to the target text information based on the associated image information; performing feature extraction on the target text information using the text feature extraction layer of the text-to-image model to obtain text feature information; performing feature extraction on the basic image information using the image feature extraction layer of the text-to-image model to obtain initial image feature information; performing sentence feature fusion on the text feature information using the feature fusion layer of the text-to-image model to obtain sentence feature information; and performing sentence feature fusion on the text feature information using the feature fusion layer of the text-to-image model to obtain sentence feature information. The feature adding layer of the text-generated image model adds random noise to the sentence feature information to obtain splicing feature information; the text image affine layer of the text-generated image model is used to perform image processing on the splicing feature information and the initial image feature information to obtain target image feature information; the convolutional network layer of the text-generated image model is used to obtain the intermediate image information corresponding to the target text information according to the target image feature information; the image enhancement layer of the text-generated image model is used to perform image refinement on the intermediate image information to obtain the target image information corresponding to the target text information; wherein, the image enhancement layer of the text-generated image model is used to perform image refinement on the intermediate image information to obtain the target image information corresponding to the target text information, including: calculating the information association value between the splicing feature information and the initial image feature through the text-image association network of the image enhancement layer; wherein, the information association value is obtained according to the following formula: ; in, represents the information association value between the i-th splicing feature information and the initial image feature, represents the i-th splicing feature information, represents the matrix required to transform the i-th splicing feature information, represents the tth initial image feature, m represents the number of features corresponding to the initial image feature, Representation conversion The required matrix, represents the activation function, and Represents a 1*1 convolution, which is used to perform dimensional conversion so that the i-th concatenated feature information and have the same feature space; An image segmentation module, configured to perform image segmentation on the target image information according to the initial image information to obtain initial object information of the target image information that does not appear in the initial image information; A virtual construction module, used to virtually construct the initial object information to obtain second three-dimensional information corresponding to the initial object information; A result generation module is used to generate mixed reality content according to the first three-dimensional information and the second three-dimensional information to obtain a mixed reality generation result corresponding to the initial voice information.

9. A terminal device, characterized in that: The terminal device includes a processor and a memory; The memory is used to store computer programs; The processor is used to execute the computer program and implement the mixed reality content generation method based on the AI ​​big model as described in any one of claims 1 to 7 when executing the computer program.

Citation Information

Patent Citations

  • Multi-modal named entity recognition method based on dependency syntax and graph neural network

    CN118673922A

  • Quantum, biological, computer vision, and neural network systems for industrial internet of things

    WO2022236064A2