Quantum type AR scene generation method based on hybrid model under Babylon.js framework

By combining CNN-HMM-QGAN hybrid model and quantum computing under the Babylon.js framework, the problems of limited acquisition equipment and poor data quality in the existing VR scene generation technology are solved, and the efficient generation of realistic and virtual scenes that meet user needs are achieved, reducing the difficulty and cost of modeling, and improving the richness and uniqueness of the generated scenes.

CN120339556AInactive Publication Date: 2025-07-18CHINA JILIANG UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510498250.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing VR scene generation technology has problems such as limited acquisition equipment, low scene authenticity, poor data quality, low data processing efficiency, high modeling cost, high content creation cost, and difficult content update and maintenance, making it difficult to quickly generate high-quality virtual scenes that meet user needs.

Method used

The CNN-HMM-QGAN hybrid model based on the Babylon.js framework is adopted, and the scene data is annotated and divided, combined with VGGNet, HMM and QGAN models for collaborative training, and complex feature fusion is used to generate realistic virtual scene data that fits the dynamic laws of the scene, and integrates and optimizes in the Babylon.js environment.

Benefits of technology

It improves the authenticity and controllability of scene generation, reduces the difficulty of modeling and content creation costs, enhances the richness and uniqueness of the generated scenes, and provides a high-quality virtual scene experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339556A_ABST
    Figure CN120339556A_ABST
Patent Text Reader

Abstract

The invention discloses a quantum type AR scene generation method based on a Babylon.js framework. The quantum type AR scene generation method comprises the steps that image sequences and depth data covering different environment conditions, time points, scene layout and various states and spatial scale information of objects are collected; key objects and scenes are accurately marked, and a training set, a verification set and a test set are randomly divided, so that data distribution of different sets is similar; a CNN model, an HMM model and a QGAN model are combined, the generation effect is improved through adversarial training, the CNN and the HMM adopt an end-to-end cooperative training mode, feedback information of the HMM serves as an additional supervision signal of the CNN, features extracted by the CNN are used for iterative optimization and mutual promotion of the HMM, the modeling ability of the models for dynamic scene change is jointly improved, and meanwhile, in the adversarial training of the QGAN, the modeling ability of the models for dynamic scene change is improved. The generator and the discriminator continuously optimize own parameters, so that the quality of the generated scene data is improved; complex feature fusion is carried out or new features are generated by using the superposition state and entanglement features of quantum calculation, and complex feature interaction is processed; through multiple rounds of iterative training and optimization, the generator generates vivid new scene data conforming to the scene dynamic law; a trained model is embedded into a Babylon.js environment, smooth data interaction is guaranteed by writing adaptive codes, generated scene elements can be completely built and optimized under the framework, and accurate building and high-quality rendering of a virtual scene are achieved through the functions of the framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of artificial intelligence, quantum methods, and virtual reality technology, and is applied to the culture and tourism industry. Specifically, it relates to a quantum-style AR scene generation method based on a CNN-HMM-QGAN hybrid model under the Babylon.js framework.

[0002] Background Method

[0003] Currently, the common scene generation methods for VR glasses include modeling and rendering methods based on computer graphics, panoramic photography and stitching methods, real-time rendering and interaction methods, and scene generation methods based on artificial intelligence. Among them, the modeling and rendering methods based on computer graphics use professional 3D modeling software, such as Blender, 3ds Max, Maya, etc., to generate realistic virtual scenes by performing operations such as geometric modeling, texture mapping, and lighting settings on the objects in the virtual scene. However, the production process is complex and cumbersome, involving a large amount of detail adjustment and optimization, resulting in high labor costs and a long production cycle. It is difficult to quickly produce a large amount of content to meet market demand; moreover, it is difficult to optimize the model, and the produced scenes also lack realistic details.

[0004] The panoramic photography and stitching method uses a 360-degree panoramic camera to take pictures of real scenes, and then stitches these pictures into a complete panoramic image through an image stitching algorithm, and then presents it to the user through the display method of VR glasses, making the user feel as if they are in the photographed real scene. However, the price of 360-degree panoramic cameras is relatively high, and the shooting quality and parameters of different cameras vary greatly, and there are also certain requirements for the shooting environment. In some special environments, the shooting effect may be seriously affected; moreover, in graphic stitching, it is necessary to accurately identify and match the feature points of adjacent photos. If there are problems such as blurring, occlusion, and repeated textures in the photos, misalignment, gaps, and blurring may occur in the stitching, affecting the quality of the panoramic image. At the same time, the calculation amount of the stitching algorithm is large, and the processing time may be very long for the stitching of a large number of photos; in addition, the panoramic image presents a fixed real scene, lacking the integration of virtual elements and creative designs, and the visual effect is relatively single.

[0005] The real-time rendering and interaction method utilizes a high-performance graphics processing unit (GPU) and a real-time rendering engine to perform real-time rendering on a virtual scene, enabling the virtual scene to respond instantaneously to the user's interaction actions, such as the movement, rotation, and scaling of objects. However, when dealing with complex virtual scenes and a large number of interaction operations, problems such as frame rate drops and lags may occur, affecting the user experience. Especially on mobile devices or low-end hardware platforms, the performance of real-time rendering is often worse. Moreover, increasing the complexity and details of the virtual scene while ensuring the performance of real-time rendering is a challenge. If the scene is too complex, it will lead to excessive rendering pressure and performance degradation. On the other hand, if too much scene detail is sacrificed for performance, it will affect the visual effect and immersion.

[0006] The artificial intelligence-based scene generation method uses artificial intelligence algorithms to learn a large amount of real-world scene data and then generates new virtual scenes. This method can quickly generate diverse virtual scenes with a certain degree of realism and creativity. However, artificial intelligence algorithms require a large amount of real-world scene data for learning and training, and the cost of data acquisition is relatively high. Moreover, the quality and diversity of the data have a great impact on the generation results. If the data has problems such as noise, bias, or incompleteness, it may lead to incorrect, unreasonable, or unrealistic virtual scenes. In addition, the controllability of the generation results is relatively poor, and it is difficult to generate precisely according to the specific needs and expectations of users. The generated scenes may contain some elements that do not conform to logic or do not meet the actual application requirements, and further manual adjustment and optimization are required. Summary of the Invention

[0007] The present invention aims to overcome the above-mentioned drawbacks of the prior art and provides a quantum-style AR scene generation method based on a CNN-HMM-QGAN hybrid model under the Babylon.js framework.

[0008] The first aspect of the present invention relates to a quantum-style AR scene generation method based on a hybrid model under the Babylon.js framework, including the following steps:

[0009] S1. Annotate the key object and scene category information for the collected scene data (image sequences and depth data containing rich visual and spatial information, ensuring that different states and detail changes of the scene are covered), and divide the training set, validation set, and test set according to a certain proportion to make the data format meet the input requirements of the subsequent model.

[0010] S2. Use VGGNet as the basic structure to construct a CNN, adjust the relevant parameters according to the characteristics of the scene, connect the output features of the CNN to the HMM, regard the feature sequence extracted by the CNN as the observation value of the HMM, and set the hidden state of the HMM based on the internal logic of the scene.

[0011] S3. Determine the initial transition probability and emission probability matrix, and use the Baum-Welch algorithm to iteratively optimize in combination with the training set to learn the dynamic change rules of the scenario. At the same time, input the training data in an end-to-end manner, enable the collaborative training of the CNN and the HMM, monitor the performance metrics of the model with the help of the validation set, and fine-tune the CNN structure and the HMM parameters until the model performance meets the requirements;

[0012] S4. Plan the number of qubits and the quantum gate sequence as needed to construct a quantum circuit, input it into the Qiskit Aer quantum simulator, simulate the state evolution of the qubits and the quantum gate operations, and observe the impact of the simulation results on the generated scenario data and perform preliminary optimization on it by adjusting the parameters and structure of the quantum circuit;

[0013] S5. Classify the scenario implicit features extracted by the CNN-HMM hybrid model. Use a classical computer to complete simple feature extraction and preliminary processing, input the preprocessed features into the quantum circuit, and use quantum computing for more complex feature fusion or generate new features with quantum characteristics;

[0014] S6. Define the network architectures of the QGAN generator and discriminator, use a small-scale dataset to verify and integrate the results obtained by the quantum computing and the results of the classical computing part. The integrated information is used as part of the input of the QGAN generator, and combined with random noise to generate "fake" scenario data;

[0015] S7. The discriminator discriminates between real and generated scenario data, and the two perform adversarial games to optimize their respective parameters. During the training stage, feed the real scenario samples and the output of the generator to the discriminator alternately every 5-10 rounds, and update the weights of the generator and the discriminator according to the discrimination results through backpropagation;

[0016] S8. After completing the training, use the validation set again to verify the improvement of the overall model performance after adding quantum computing and QGAN adversarial training. After multiple rounds of iteration, prompt the generator to output new scenario data that is realistic and conforms to the dynamic laws of the scenario;

[0017] S9. Embed the trained CNN-HMM hybrid model and QGAN model into the Babylon.js environment, ensure smooth data interaction according to the adaptation code, enable Babylon.js to retrieve the generation or processing results of the model to participate in the scenario construction process, and generate tests based on this to check the functions of the integrated model and framework;

[0018] S10. Inside the framework, combine the scene elements generated by the QGAN with the CNN-HMM analysis results, and use the Babylon.js framework to completely build the virtual scene. Precisely set the object coordinates, angles, and hierarchical relationships. For display requirements, optimize the rendering parameters, frame rate, and resolution, enable anti-aliasing and shadow optimization technologies to optimize the picture and light and shadow visual effects, and restore the real scene and reflect the predicted changes.

[0019] Preferably, the annotation and division of the scene data in step S1 include:

[0020] Collect scene image sequences and depth data under different environmental conditions, time points, and scene layouts. The data covers various states of the scene, including the scene under different light intensities, different placement positions and angles of objects, dynamic changes of the objects in the scene, and different spatial scale information, ensuring that the diversity and complexity of the real scene can be comprehensively reflected;

[0021] Annotate the collected data. For key objects in the image sequence, use bounding box or pixel-level annotation methods to accurately mark the position and contour of each key object and assign its corresponding class label. At the same time, also perform class annotation on the entire scene;

[0022] Randomly divide the annotated data set into a training set, a validation set, and a test set according to a certain proportion. During the division process, ensure that the data distributions in different sets are as similar as possible to avoid adverse effects of data deviation on model training and evaluation. At the same time, perform format conversion and preprocessing on the data to make it meet the input requirements of subsequent models such as CNN and HMM.

[0023] Preferably, the construction of the CNN under the VGGNet structure and the integration initialization of the HMM in step S2 include:

[0024] Build the CNN with VGGNet as the basic structure. First, determine the number of layers of the network and appropriately adjust the number of layers of the classic structure of VGGNet according to the complexity of the scene data;

[0025] Determine the ReLU function as the activation function to effectively solve the gradient vanishing problem and accelerate the training speed of the network. At the same time, configure the pooling layer to reduce the resolution of the feature map, reduce the amount of data, and extract the main features. In the last few layers of the network, set the fully connected layer according to the task requirements. The number of neurons in the fully connected layer can be adjusted according to the number of scene categories or feature dimensions, and is used to map the features extracted by the convolutional layer to the final output category or feature representation;

[0026] Input the features output by the CNN into the HMM to determine the observations of the HMM, that is, use the feature sequence extracted by the CNN as the observation sequence of the HMM, and set the hidden states of the HMM based on the internal logic of the scenario, the state changes of the objects in the scenario, factors such as the transition of the scenario, etc. The hidden states reflect the dynamic change rules behind the scenario and are associated with the feature sequence extracted by the CNN, laying a foundation for subsequent learning of the scenario dynamic change rules.

[0027] Preferably, the CNN-HMM collaborative training in step S3 includes:

[0028] Determine the initial transition probability matrix and emission probability matrix of the HMM. The transition probability matrix represents the transition probability between hidden states, and the emission probability matrix represents the probability of observing a specific feature in each hidden state. Initially, set the probability values according to prior knowledge or a simple uniform distribution assumption, and these initial values will be continuously updated in subsequent iterative optimizations;

[0029] Use the Baum-Welch algorithm to iteratively optimize the parameters of the HMM in combination with the training set. In each iteration, first calculate the probability distribution of the observation sequence in each hidden state according to the current model parameters (the transition probability matrix and the emission probability matrix) (E step), and then re-estimate the model parameters according to the probability distribution to maximize the likelihood function (M step). During this process, input the training data in an end-to-end manner to enable the CNN and the HMM to be collaboratively trained;

[0030] When training the CNN, use the feedback information of the HMM as an additional supervision signal to adjust the weights of the CNN, so that the features extracted by the CNN are more conducive to the HMM's modeling of the scenario dynamic changes. Conversely, during the iterative optimization process of the HMM, use the more accurate feature sequence extracted by the CNN to continuously update the transition probability matrix and the emission probability matrix;

[0031] After completing a certain number of rounds of the training, use the validation set to verify the model, calculate the performance metrics of the model on the validation set, and according to the verification results, fine-tune the structural parameters of the CNN and the parameters of the HMM, and repeat the training and verification process until the model performance meets the preliminary requirements.

[0032] Preferably, the quantum circuit construction and simulation optimization in step S4 includes:

[0033] Determine the number of qubits according to the complexity of the scene data and the dimension and characteristics of the features extracted by the CNN-HMM hybrid model. When making a preliminary estimate, start with a relatively small number of qubits and then gradually increase the number of qubits based on subsequent simulation results and the evaluation of the accuracy of the feature representation;

[0034] Construct a quantum gate sequence based on the basic gate operations of quantum computing to achieve specific transformations and processing of the input features, achieving the effect of complex feature fusion or generating new features. First, use single-bit gates to initialize the qubits and set them to specific initial states, and then use multi-bit gates to achieve entanglement and interaction between the qubits, thereby simulating the complex correlations between the features;

[0035] Input the constructed quantum circuit into the Qiskit Aer quantum simulator. In the simulator, set the simulation parameters. After starting the simulation, observe the state evolution process of the qubits, that is, as the quantum gate operations are sequentially executed, the process in which the qubits gradually change from the initial state to the final state, and analyze the influence of the quantum gate operations on the qubit states;

[0036] According to the simulation results, adjust the parameters and structure of the quantum circuit, change parameters such as the operation angles of the quantum gates or the action order of the gates, try to add or delete certain quantum gates, or change the connection method between the quantum gates to adjust the structure. By continuously adjusting and re-simulating, optimize the quantum circuit so that it can better process the scene features and improve the ability to generate new features with quantum characteristics and meeting the scene requirements.

[0037] Preferably, the feature classification and quantum processing in step S5 include:

[0038] Classify the scene implicit features extracted by the CNN-HMM hybrid model according to the type of the features, the importance of the features, or the correlation between the features and different aspects of the scene. Classify some simple and low-level features into one category, and these features can be preliminarily processed through simple mathematical operations on a classical computer, while classify some features related to the overall structure of the scene or the complex relationships between objects into another category, and these features will be input into the quantum circuit for more complex processing;

[0039] For simple feature extraction and preliminary processing, perform operations on a classical computer, then encode and convert the features after classical preprocessing according to the input requirements of the quantum circuit, and convert the classical data into the representation form of quantum states. Map the feature vector to the states of the qubits through a specific encoding algorithm to ensure that the quantum circuit can correctly receive and process this feature information;

[0040] Input the converted features into the quantum circuit for processing. The quantum circuit utilizes the superposition state and entanglement characteristics of the qubits, and performs complex feature fusion operations on the input features through the quantum gates, enabling entanglement to occur between the features encoded by different qubits, thereby simulating complex feature interactions that are difficult to achieve in classical computing.

[0041] Preferably, the determination and data integration of the QGAN architecture in step S6 include:

[0042] Use a small-scale dataset to verify and integrate the results obtained from the quantum computing and the results of the classical computing part. First, run the codes of the quantum computing and the classical computing parts respectively on the small-scale dataset to obtain their respective output results. Then, evaluate the quality and effectiveness of the results through some verification metrics;

[0043] According to the verification results, integrate the results of the quantum computing and the classical computing. If it is found that some of the quantum computing results are insufficient in certain aspects, make appropriate adjustments or corrections to them, and use the integrated information as part of the input of the QGAN generator, and combine random noise to generate "fake" scenario data. The addition of the random noise can increase the diversity and randomness of the generated data, enabling the generator to explore a wider scenario data space and avoid generating overly single or deterministic scenario data;

[0044] During the generation process, ensure that the integrated input information and the random noise can be reasonably processed and fused in the network of the generator to generate high-quality "fake" scenario data, laying a foundation for subsequent QGAN training and scenario construction.

[0045] Preferably, the QGAN adversarial training in step 7 includes:

[0046] Before starting the training, first initialize the network parameters of the generator and discriminator of the QGAN. The generator parameters are randomly initialized with relatively small random values to avoid excessive gradient updates at the beginning of the training. The discriminator parameters are also initialized similarly. At the same time, prepare a representative real scenario sample dataset, which needs to cover various scenario states and types and match the output of the generator in terms of data format and feature representation for effective comparison and discrimination;

[0047] Enter the training loop. Data feeding and parameter updating are performed every 5 - 10 rounds. In each round, first, the generator generates a batch of "fake" scenario data based on its current parameters and input. Then, real scenario samples and the generator output are alternately fed to the discriminator. The discriminator discriminates these data and outputs the probability estimates of whether each data sample is real scenario data or generated data.

[0048] Calculate the loss function value of the discriminator according to the discrimination results. The binary cross - entropy loss function is adopted. For real samples, the discriminator outputs a probability close to 1, and for generated samples, it outputs a probability close to 0. Calculate the gradient of the discriminator loss function with respect to its parameters through the backpropagation algorithm, and use the Adam optimizer to update the discriminator parameters, so that the discriminator can better distinguish real and generated data.

[0049] Next, use the generator to generate another batch of "fake" scenario data and input it into the discriminator with the updated parameters. This time, calculate the loss function of the generator to make the discrimination result of the discriminator for the generated data close to 1, that is, the generated data can "fool" the discriminator. Similarly, calculate the gradient of the generator loss function with respect to its parameters through backpropagation and update the generator parameters. In this way, the generator and the discriminator continuously optimize their own parameters during the adversarial process and gradually improve the ability of the generator to generate realistic scenario data.

[0050] Preferably, the model performance verification and iterative optimization in step S8 include:

[0051] After completing a certain number of rounds of adversarial training, prepare a validation set. The validation set is independent and similar to the training set and the test set in data distribution, and contains samples of different scenario categories and states. Determine the metrics for evaluating the model performance, including visual quality metrics for scenario generation, degree of fit metrics for scenario dynamic laws, and other metrics related to specific application scenarios.

[0052] Input the validation set data into the overall model after the quantum computing and QGAN adversarial training to obtain the generated scenario data. Evaluate the generated data using the pre - set metrics, calculate the values of each metric and compare them with the results before training or in the previous validation rounds to analyze the improvement of the overall model performance after adding the quantum computing and QGAN adversarial training.

[0053] According to the verification results, adjust and optimize the model. If the number of training rounds is insufficient, increase the number of training rounds; if there are integration issues, re-examine the data interaction and parameter transfer methods between the quantum computing and the QGAN, and make appropriate modifications; if there are model architecture issues, try to adjust the network structures of the QGAN generator or the discriminator by increasing or decreasing the number of layers, changing the neuron connection methods, etc., and then perform training and verification again. After multiple rounds of iteration, until the generator can produce new scene data that is realistic and conforms to the dynamic laws of the scene, meeting the application requirements.

[0054] Preferably, the model embedding and integration testing in step S9 includes:

[0055] Integrate the trained CNN-HMM hybrid model and the QGAN model into the Babylon.js environment. First, write adaptation code to ensure that the code structure and data interfaces of the model are compatible with the Babylon.js framework. At the same time, handle the function call relationship between the model and Babylon.js so that Babylon.js can call the model at the appropriate time for scene generation or processing operations;

[0056] Conduct some simple scene element generation tests to check the functions after the model is integrated with the framework. Observe whether the elements generated by the model can be correctly displayed in the Babylon.js scene, check whether their attributes such as position, size, and direction meet the expectations, verify whether the data interaction between the model and the framework is smooth, and whether the parameters passed from Babylon.js to the model can be correctly received and processed by the model. At the same time, check whether the model can work properly in the new Babylon.js environment to ensure the stability and reliability of the entire integrated system.

[0057] Preferably, the final scene construction and optimization in step S10 includes:

[0058] Within the Babylon.js framework, complete the virtual scene construction based on the scene elements generated by the QGAN and combined with the CNN-HMM analysis results. Use various scene objects, textures, lighting effects, etc. generated by the QGAN, and according to the CNN-HMM analysis of the scene structure and dynamic laws, determine the positions, angles, and hierarchical relationships of these elements in the scene;

[0059] For display requirements, optimize rendering parameters, frame rate, resolution, etc. According to the requirements of the target platform and application scenario, select an appropriate resolution, adjust the frame rate to ensure the smoothness of the animation and interaction in the scene, optimize the rendering parameters, enable anti-aliasing technology to reduce the jagged effect at the edges of the image, and enable shadow optimization technology to make the shadows in the scene more realistic and natural, so as to comprehensively optimize the picture and light and shadow visual effects, restore the real scene and reflect the predicted changes, and provide users with a high-quality virtual scene experience.

[0060] The advantages of the present invention are:

[0061] Traditional VR scene generation technologies more or less face problems such as limited acquisition devices, low scene authenticity, poor data quality, low data processing efficiency, high modeling difficulty, high content creation cost, and difficulties in content updating and maintenance. In response to this, the present invention can comprehensively reflect the diversity and complexity of real scenes by collecting image sequences and depth data covering different environmental conditions, time points, scene layouts, and various states and spatial scales of objects, providing rich materials for subsequent model training and making the generated scenes closer to reality; by accurately annotating key objects and scenes, randomly dividing the training set, validation set, and test set according to a certain ratio, and ensuring that the data distributions of different sets are similar, avoiding the adverse effects of data deviation on model training and evaluation, and improving the generalization ability and reliability of the model; combining the advantages of three models, namely CNN, HMM, and QGAN. CNN has a powerful feature extraction ability and can automatically learn the features in images; HMM can model the dynamic changes of scenes and learn the logical rules behind the scenes. QGAN can generate high-quality "fake" scene data, and continuously improve the generation effect through adversarial training. The three work together to make scene generation more accurate and realistic; CNN and HMM adopt an end-to-end collaborative training method. The feedback information of HMM is used as an additional supervision signal for CNN, and the features extracted by CNN are used for the iterative optimization of HMM, promoting each other and jointly improving the model's ability to model the dynamic changes of scenes. At the same time, in the adversarial training of QGAN, the generator and discriminator continuously optimize their own parameters, further improving the quality of the generated scene data; using the superposition state and entanglement characteristics of quantum computing for complex feature fusion or generating new features, which can handle some complex feature interactions that are difficult to achieve in classical computing, bringing more possibilities and innovation to scene generation, and enhancing the richness and uniqueness of the generated scenes; after multiple rounds of iterative training and optimization, the generator can output new scene data that is realistic and conforms to the dynamic laws of the scene, performing well in terms of visual quality and the degree of conformity with the dynamic laws of the scene, and being able to better meet the user's requirements for realism and logic; embedding the trained model into the Babylon.js environment, ensuring smooth data interaction by writing adaptation code, enabling the generated scene elements to be fully built and optimized under this framework, and using the functions of the framework to accurately construct and high-quality render virtual scenes, providing users with a good visual experience. Brief Description of the Drawings

[0062] Figure 1 The figure shows an overview diagram of the system of the quantum AR scene generation method based on the CNN-HMM-QGAN hybrid model under the Babylon.js framework in a specific embodiment of the present invention.

[0063] Figure 2The figure shows the system flow chart of the quantum AR scene generation method based on the CNN-HMM-QGAN hybrid model under the Babylon.js framework in the specific embodiment of the present invention. Detailed implementation manners

[0064] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0065] This embodiment relates to a quantum AR scene generation method based on a hybrid model under the Babylon.js framework, which is characterized by including the following steps:

[0066] S1. Label the key object and scene category information for the collected scene data (image sequences and depth data containing rich visual and spatial information, ensuring that different states and detail changes of the scene are covered), divide the training set, validation set, and test set according to a certain proportion, and make the data format meet the input requirements of the subsequent model;

[0067] S2. Use VGGNet as the basic structure to construct the CNN, adjust the relevant parameters according to the scene characteristics, connect the output features of the CNN to the HMM, regard the feature sequence extracted by the CNN as the observation value of the HMM, and set the hidden state of the HMM based on the internal logic of the scene;

[0068] S3. Determine the initial transition probability and emission probability matrix, use the Baum-Welch algorithm to iteratively optimize in combination with the training set, learn the dynamic change rules of the scene, and at the same time input the training data in an end-to-end manner to co-train the CNN and the HMM. Monitor the performance indicators of the model with the validation set, and fine-tune the CNN structure and the HMM parameters until the model performance meets the requirements;

[0069] S4. Plan the number of qubits and the quantum gate sequence as needed to construct a quantum circuit, input it into the Qiskit Aer quantum simulator, simulate the state evolution of the qubits and the quantum gate operations, and observe the influence of the simulation results on the generated scene data and perform preliminary optimization on it;

[0070] S5. Classify the scene implicit features extracted by the CNN-HMM hybrid model. Use a classical computer to complete simple feature extraction and preliminary processing, input the preprocessed features into the quantum circuit, and use quantum computing for more complex feature fusion or generate new features with quantum characteristics;

[0071] S6. Define the network architectures of the QGAN generator and discriminator. Use a small-scale dataset to verify and integrate the results obtained from quantum computing with those of the classical computing part. The integrated information serves as part of the input to the QGAN generator, and combined with random noise, it generates "fake" scenario data.

[0072] S7. The discriminator distinguishes between real and generated scenario data. The two engage in an adversarial game to optimize their respective parameters. During the training phase, every 5 - 10 rounds, alternately feed the real scenario samples and the output of the generator to the discriminator. According to the discrimination results, backpropagate to update the weights of the generator and discriminator.

[0073] S8. After completing the training, use the validation set again to verify the improvement in the overall model performance after adding quantum computing and QGAN adversarial training. After multiple rounds of iteration, prompt the generator to produce new scenario data that is realistic and conforms to the dynamic laws of the scenario.

[0074] S9. Embed the trained CNN - HMM hybrid model and QGAN model into the Babylon.js environment. Ensure smooth data interaction according to the adaptation code, so that Babylon.js can retrieve the generated or processed results of the model to participate in the scenario construction process, and use this to generate tests to check the functions after the integration of the model and the framework.

[0075] S10. Within the framework, combine the scene elements generated by the QGAN with the analysis results of the CNN - HMM. Use the Babylon.js framework to completely build the virtual scene, accurately set the object coordinates, angles, and hierarchical relationships. For display requirements, optimize the rendering parameters, frame rate, and resolution, and enable anti - aliasing and shadow optimization techniques to optimize the picture and light and shadow visual effects, restoring the real - world scene and reflecting the predicted changes.

[0076] Use a high - resolution camera or camera to capture various objects and details in the scene. According to the characteristics and requirements of the scene, determine appropriate shooting angles, distances, and lighting conditions to obtain a representative image sequence. Use an Intel RealSense depth sensor to work synchronously with the image acquisition device to obtain the depth information of the objects in the scene. At the same time, pay attention to the accuracy and integrity of the depth data during the acquisition process to avoid data loss or excessive noise.

[0077] Use the VGG Image Annotator image annotation tool to box and classify the key objects and depth data in the images. It also supports relevant annotation operations. For key objects, determine the annotation methods for their specific names, positions, sizes, etc.; for scene categories, classify them according to the functions, environments, and other characteristics of the scene, and formulate corresponding annotation standards.

[0078] Using the method of random sampling, divide the labeled data into a training set, a validation set, and a test set according to a certain ratio, ensuring that the data in each set is representative and independent, and avoiding problems such as data leakage and overfitting.

[0079] According to the input requirements of the subsequent model, perform format conversion on the divided data sets, convert the image data into the tensor format required by the model, convert the annotation information into the corresponding encoding format, etc. At the same time, perform preprocessing operations such as normalization and standardization on the data to improve the training effect and generalization ability of the model.

[0080] According to the characteristics of the scene data and the task requirements, select the appropriate version of VGGNet. Since different versions vary in terms of the number of network layers, the number of parameters, etc., factors such as computing resources, training time, and model performance should be comprehensively considered when making a selection.

[0081] Analyze the characteristics such as the size, shape, and texture of the objects in the scene data, and select the appropriate convolutional kernel size. For scenes containing larger objects, appropriately increase the convolutional kernel size to better capture the overall features of the objects; for scenes with rich details, increase the number of convolutional kernels to extract more local features.

[0082] According to the resolution and object distribution of the scene data, adjust the parameters such as the pooling kernel size and stride of the pooling layer. For high-resolution scene data, appropriately increase the pooling kernel size and stride to reduce the data dimension while retaining the key features.

[0083] According to the number of scene categories and the classification difficulty, adjust the number of neurons and activation functions in the fully connected layer. If there are more scene categories and the classification difficulty is greater, increase the number of neurons in the fully connected layer and select a more complex activation function to improve the non-linear expression ability of the model.

[0084] Use the TensorFlow deep learning framework to build a CNN model based on VGGNet, set the network components such as the convolutional layer, pooling layer, and fully connected layer according to the adjusted parameters, and define the forward propagation process to ensure that the model can correctly extract features and classify the input data.

[0085] Connect the output features of the CNN to a Hidden Markov Model (HMM). Use the feature sequence extracted by the CNN as the observations of the HMM, and match the output dimension of the CNN with the observation space dimension of the HMM through a fully connected layer or other means to ensure the smooth transfer and processing of the data.

[0086] Based on the internal logic and dynamic change characteristics of the scene, set the hidden states of the HMM so that the model can learn the dynamic change rules of the scene.

[0087] Initialize the transition probabilities between the hidden states of the HMM according to the prior knowledge and experience of the scenario, and set an initial transition probability matrix according to the scenario logic; similarly, based on the prior knowledge, determine the emission probabilities from the hidden states to the observations (i.e., the feature sequences extracted by the CNN), so as to initialize the emission probability matrix.

[0088] Input the divided training set data into the CNN-HMM model, where the CNN extracts features from the image sequence to obtain the feature sequence as the observation of the HMM.

[0089] In each iteration, use the Baum-Welch algorithm to update the transition probability matrix and the emission probability matrix. Through the statistical analysis of the training data, re-estimate the transition probabilities between the hidden states and the emission probabilities from the hidden states to the observations, so that the model can better fit the training data.

[0090] The specific calculation process is as follows: for each observation sequence in the training set, use the forward-backward algorithm to calculate the forward probability and backward probability under the current model parameters, and then obtain the posterior probability of each hidden state at each time. According to these posterior probabilities, update the transition probability matrix and the emission probability matrix. Among them, the update formula for the transition probability is: where a ij is the transition probability from the hidden state i to the hidden state j, ξ t (i,j) is the probability of being in the hidden state i at time t and in the hidden state j at time t+1, γ t (i) is the probability of being in the hidden state i at time t; the update formula for the emission probability is: where b i (k) is the probability that the hidden state j emits the observation value v k , and o t is the observation value at time t. Continuously repeat the above process until the model parameters converge, that is, the changes in the transition probability matrix and the emission probability matrix are less than a certain threshold, or the preset number of iterations is reached.

[0091] Input the training data into the CNN and HMM in an end-to-end manner at the same time, and let the CNN and HMM be trained simultaneously. During the training process, the parameters of the CNN are updated based on its own cross-entropy loss function, which is used to measure the accuracy of the CNN's prediction of object and scene categories; the parameters of the HMM are updated based on the optimization of the transition probabilities and emission probabilities by the Baum-Welch algorithm. At the same time, the CNN and HMM interact through feature transfer. The features output by the CNN are used as the observations of the HMM, and the hidden state information of the HMM can also be fed back to the CNN to help the CNN better learn the scene features.

[0092] During the training process, a validation set is used to monitor the performance metrics of the model, such as accuracy, recall, F1-score, etc., which are used to evaluate the accuracy of the CNN's prediction of object and scene categories, and perplexity, which is used to measure the goodness of fit of the HMM to the observation sequence.

[0093] According to the monitored metric performance, fine-tune the structure of the CNN and the parameters of the HMM. If the accuracy is found to be low, the network structure of the CNN can be adjusted, such as adding or reducing convolutional layers, adjusting the size of the convolutional kernel, etc.; if the perplexity is high, further optimize the transition probability and emission probability matrices of the HMM, or adjust the feature extraction method of the CNN to improve the model's learning ability for scene dynamic changes until the model performance meets the requirements.

[0094] Analyze the dimensions, complexity of the features extracted by the CNN-HMM hybrid model, and the characteristics of the scene data itself. If the scene data involves many different types of objects, complex spatial relationships, and dynamic changes, the corresponding extracted feature dimensions will be higher and the correlations will be more complex. Then, initially, a relatively larger number of qubits can be considered; conversely, if the scene is relatively simple and the features are relatively single, then start with a smaller number of qubits.

[0095] Start with 2 - 4 qubits, input the constructed quantum circuit into a quantum simulator for simulation, observe the effect of its feature processing and whether it can effectively represent the scene features. According to the simulation results, if it is found that the complex relationships between features cannot be fully reflected or new features that meet the requirements cannot be well generated, gradually increase the number of qubits, increasing by 1 - 2 each time, and perform simulation and evaluation again until it is considered that the number of qubits can better adapt to the requirements of scene feature processing.

[0096] Use the Hadamard single-qubit gate to initialize the qubits, set them to a specific initial state, prepare the qubits into a superposition state, enabling them to represent multiple possibilities simultaneously, facilitating the subsequent simulation of complex correlations between features and laying the foundation for more complex operations later.

[0097] Based on the basic principles of quantum computing and the desired feature transformation effect, use the CNOT multi-qubit gate to achieve entanglement and interaction between qubits, so that the state of the target qubit will change according to the state of the control qubit, thereby simulating the relationship of mutual influence and mutual restriction between different feature elements in the real scene, and thus realizing functions such as complex feature fusion, and constructing a quantum gate sequence that meets the requirements of specific transformation and processing of the input features.

[0098] Input the constructed quantum circuit into the Qiskit Aer quantum simulator, and then set relevant simulation parameters, including but not limited to the number of simulation time steps (corresponding to the number of executions of quantum gate operations), measurement methods, initial state parameters, etc., to ensure that the simulation proceeds accurately as expected and meets the requirements for observing the state evolution of qubits and quantum gate operations.

[0099] After starting the simulation process, carefully observe the state evolution of the qubits, that is, the entire process in which the qubits gradually change from the initially set initial state to the final state as the quantum gate operations are executed sequentially. Pay attention to the changes in the qubit states after each quantum gate operation and analyze how these changes affect the representation and processing effects of the scene features.

[0100] Based on the analysis of the simulation results, try to adjust the parameters of the quantum circuit, change the operation angles of the quantum gates. Different angles will cause different transformations of the qubit states, thus affecting the feature fusion effect; or adjust the order of the gate actions. Changing the order of quantum gate operations often results in different evolution paths of the qubit states and different final states, so as to observe whether the quantum circuit can better process the scene features and get closer to the goal of generating new features that meet the scene requirements.

[0101] In addition to parameter adjustment, you can also try to change the structure of the quantum circuit, such as adding or deleting certain quantum gates, which may introduce new interaction relationships between qubits or remove some unnecessary operations; or change the connection method between quantum gates and re-plan the entanglement layout between qubits. By continuously adjusting the structure and re-performing the simulation, find the optimal quantum circuit structure and improve its ability to generate new features with quantum characteristics and meet the scene requirements.

[0102] Consider the scene implicit features extracted by the CNN-HMM hybrid model from multiple dimensions, including the type of features (whether they are geometric features such as describing the shape, color, position of objects, or time series features reflecting the dynamic changes of the scene, etc.), the importance of the features for the overall understanding of the scene (key dominant features or relatively minor auxiliary features), and the correlation of the features with different aspects of the scene (the degree of association with object movement, scene layout, lighting conditions, etc.).

[0103] Based on the above analysis, specific classification criteria are formulated to classify the features. The features that directly reflect the basic attributes of the object's appearance and whose information can be further refined through simple mathematical operations are grouped into one category. Such features are usually intuitive and simple. The features that involve the complex spatial relationships and dynamic change laws among multiple objects in the scene and have deep logical connections with each other are grouped into another category. Such features often require more complex processing mechanisms to extract the value they contain, and will be input into the quantum circuit for processing later.

[0104] For the category of simple features, common mathematical operations, statistical methods, or basic image processing techniques are used on a classical computer for processing. If it is a feature describing the position of an object, statistical quantities such as the average position at different times and the standard deviation of position changes can be calculated; if it is a color feature, operations such as color space conversion and color histogram statistics can be performed to initially refine and regularize these features to make them more regular and representative.

[0105] After completing the simple feature extraction and processing, encoding and conversion operations are performed on the data according to the input requirements of the quantum circuit. Based on specific encoding algorithms in quantum computing, the feature data represented in the form of conventional numbers, vectors, etc. on a classical computer is converted into the representation form of quantum states to ensure that the subsequent quantum circuit can correctly receive and process this feature information. Common encoding methods include amplitude encoding, angle encoding, etc. According to the characteristics of the features and the design requirements of the quantum circuit, a suitable encoding method is selected to accurately map the feature vector to the state of quantum bits.

[0106] The feature data that has been preprocessed classically and converted into the representation form of quantum states is input into the constructed quantum circuit. Utilizing the superposition state and entanglement characteristics of quantum bits, through the operations of various quantum gates in the quantum circuit, complex feature fusion processing is performed on the input features to simulate the complex feature interaction effects that are difficult to achieve in classical computing, enabling the originally relatively independent features in the classical space to have deep associations in the quantum space, and then generating new feature representations with quantum characteristics, and extracting more information hidden in the scene data, laying a foundation for generating high-quality scene data later.

[0107] After being processed by the quantum circuit, the new features with quantum characteristics will be used as important inputs for subsequent steps. According to the established process and data interface requirements, these output features are passed to relevant model components such as QGAN in the subsequent steps to ensure the data coherence and coordination of the entire scene generation technology process, so as to further achieve the scene generation task based on quantum characteristics.

[0108] Carefully select a representative small-scale dataset that covers various typical situations in the scenario, including different object combinations, diverse spatial layouts, and various dynamic changes, etc., and is consistent with the subsequent actual data in terms of data format, feature distribution, etc., so that it can effectively reflect the characteristics of the real scenario and the processing effect of the model.

[0109] On this small-scale dataset, run the codes of the quantum computing part and the classical computing part respectively, ensuring that each part operates according to the established model architecture and parameter settings. Let the quantum computing part process the input features based on the quantum circuit constructed and optimized in the previous steps and output the corresponding quantum computing results; the classical computing part outputs the corresponding results according to its corresponding processing logic through the classical operation process in the CNN-HMM hybrid model.

[0110] According to the task requirements and data characteristics, determine appropriate verification metrics to evaluate the quality and effectiveness of the quantum computing results and the classical computing results. Commonly used metrics include the mean squared error (MSE), which is used to measure the average degree of difference between the results and the true values (if there are reference annotated true values) or expected values; the cosine similarity, which is used to evaluate the similarity in direction between two result vectors to judge the similarity of feature representations; and the structural similarity index (SSIM), etc., which is especially suitable for image-like scenario data and can measure the similarity between the two from multiple aspects such as brightness, contrast, and structure, etc.

[0111] Through the selected verification metrics, carefully compare the differences between the quantum computing results and the classical computing results, and observe in which aspects the quantum computing shows advantages; at the same time, pay attention to the possible deficiencies in the quantum computing results and in which dimensions the classical computing results can provide reliable and complementary information.

[0112] According to the situation of the comparative analysis, make appropriate adjustments or corrections to the quantum computing results. If it is found that the numerical range of certain features in the quantum computing results is unreasonable or there is a large conflict with the classical computing results, based on the reasonable parts of the classical computing part and the understanding of the scenario logic, adjust the quantum computing results by setting reasonable thresholds, performing feature normalization, etc., so that it better meets the actual scenario requirements and can be better integrated with the classical computing results.

[0113] Integrate the adjusted and corrected quantum computing results and the classical computing results according to the predetermined rules. Weighted sum the corresponding feature vectors of the two according to a certain weight, or combine them into a new feature representation form by splicing, etc., to ensure that the integrated information can completely and reasonably reflect the scenario features, while retaining the unique advantages brought by quantum computing and the reliability of classical computing.

[0114] The integrated information is used as part of the input to the QGAN (Quantum Generative Adversarial Network) generator. At the same time, to increase the diversity and randomness of the generated data, an appropriate random noise is added to the generator. The characteristics such as the dimension and distribution of this random noise need to be reasonably designed according to the network architecture of the QGAN and the characteristics of the expected generated scenario data. A random noise vector with a normal distribution is used, and its dimension matches that of the integrated input information so that it can be smoothly fused in the generator network.

[0115] Inside the QGAN generator, through its designed network architecture (such as including multiple convolutional layers, fully connected layers, etc., determined according to specific applications), complex non-linear transformation and mapping operations are performed on the input integrated information and random noise, enabling the generator to generate "fake" scenario data that meets the requirements based on these inputs. During this process, ensure that the parameters of each layer in the network can be reasonably adjusted according to the input information to achieve an effective feature transformation from input to output, and finally generate "fake" scenario data with a certain degree of innovation, diversity, and in line with the scenario logic, providing basic materials for subsequent adversarial training and scenario construction.

[0116] For the generator network of the QGAN, its parameters are randomly initialized with small random values. Small random numbers following a uniform distribution or a normal distribution are used to initialize the weights of the convolutional layers, the weights and biases of the fully connected layers, etc., to avoid excessive gradient updates at the beginning of training, resulting in an unstable training process and difficulty in convergence.

[0117] Similarly, the parameters of the discriminator network are initialized in a similar way so that it can participate in the subsequent training in a relatively reasonable initial state. The network structure of the discriminator usually also consists of convolutional layers, fully connected layers, etc. According to its design architecture, appropriate small random values are assigned to the weights and biases of each layer to initialize the parameters.

[0118] Collect a dataset of real scenario samples covering various scenario states and types, and ensure that this dataset matches the output of the generator in terms of data format and feature representation for subsequent effective comparison and discrimination operations. Necessary preprocessing, such as normalization, is performed on the collected data to make it within a suitable data range for model processing.

[0119] Determine the number of rounds required for the entire training process. The specific number of rounds is determined according to factors such as the convergence of the model and the quality requirements of the generated scenario data. At the same time, it is stipulated that a data feeding and parameter update operation between the real scenario samples and the output of the generator is performed every 5 - 10 rounds to balance the training efficiency and the model convergence effect to a certain extent, and to avoid the impact of overly frequent or sparse data feeding on the training quality.

[0120] In each round of training, first, the generator generates a batch of "fake" scenario data according to its current parameter settings and input information (including the integrated quantum computing and classical computing results and added random noise). The generator gradually transforms the input information into output data with certain visual features and scenario structures through its internal network structure, simulating the features and distributions of real scenario data.

[0121] The real scenario samples and the "fake" scenario data output by the generator are fed to the discriminator alternately. After receiving these data, the discriminator uses its own network structure (usually including multiple convolutional layers for feature extraction and fully connected layers for classification judgment, etc.) to process each data sample and outputs a probability estimate of whether each data sample is real scenario data or generated data. For real scenario samples, the discriminator outputs a probability value close to 1 after calculation, indicating that it believes this sample is probably real scenario data; while for the "fake" scenario data output by the generator, it outputs a probability value close to 0, indicating that it judges this sample as generated "fake" data.

[0122] The binary cross-entropy loss function is used to measure the discrimination effect of the discriminator. This loss function calculates the loss value based on the difference between the probability estimates of the real and generated samples output by the discriminator and the actual labels (the real sample label is 1, and the generated sample label is 0). Its calculation formula is roughly as follows: where m is the number of samples, y i is the real label of sample i, is the probability estimate of the discriminator's output for the sample being real scenario data. By calculating the loss function value of the discriminator through this formula, it reflects the gap between the current discrimination ability of the discriminator and the ideal situation.

[0123] Using the backpropagation algorithm, the gradient of the discriminator loss function with respect to its parameters is calculated based on the obtained loss function value, and the gradient value corresponding to each parameter is obtained. This gradient value represents the direction and magnitude of parameter adjustment. Then, the Adam optimizer is used to update the parameters of the discriminator according to the gradient value, adjusting the weights and biases and other parameters in its network, so that the discriminator can better distinguish real and generated data, optimize in the direction of reducing the loss function value, and improve the discrimination ability.

[0124] After the parameters of the discriminator are updated, the generator generates another batch of "fake" scenario data again and inputs it into the discriminator with updated parameters. At this time, the loss function of the generator is calculated. The goal is to make the discriminator's discrimination result for the generated data close to 1, that is, it is hoped that the generated data can "deceive" the discriminator and make it difficult to distinguish between true and false. The loss function of the generator can also be constructed based on a suitable loss function such as binary cross-entropy. Its core idea is to make the generated data be judged as real data as much as possible in the eyes of the discriminator, so as to measure the ability of the generator to generate realistic scenario data.

[0125] Calculate the gradients of the generator loss function with respect to its own parameters through the backpropagation algorithm to obtain the direction and magnitude information of the parameters that need to be adjusted. Then, use optimizers such as Adam to update the parameters of the generator based on these gradient information, and optimize the network structure of the generator so that it can generate "fake" scene data with higher quality and more realism.

[0126] Select a validation set from the labeled and partitioned dataset to ensure that the validation set is independent and similar to the training set and test set in terms of data distribution. That is, the validation set should contain samples of different scene categories and cover various scene states, so that it can comprehensively and objectively reflect the performance of the model in different actual scenarios, and avoid inaccurate evaluation of the model performance due to data bias.

[0127] Determine the appropriate number of samples and proportion of the validation set according to the size of the overall dataset. If the overall dataset is large, 10%-20% of the corresponding number of samples can be selected to form the validation set; if the dataset itself is small, under the premise of ensuring the validation effectiveness, as many representative samples as possible should be selected to ensure that the model performance can be fully verified.

[0128] Adopt metrics such as peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) to measure the image clarity of the generated scene data and the similarity with the real scene in terms of structure, brightness, contrast, etc. PSNR reflects the image quality by comparing the differences of corresponding pixel points between the generated image and the real image. The higher the value, the clearer the image and the smaller the distortion; SSIM comprehensively evaluates the similarity of the image from multiple dimensions, and the value range is between 0 and 1. The closer to 1, the higher the visual similarity with the real scene.

[0129] Measure the detail performance by counting the number of detail features of objects in the generated scene, the richness of textures, etc., and calculate the sharpness of the edges of objects and the complexity of textures in the scene, etc., to judge whether the generated scene can present rich detail information like the real scene and avoid blurring and lack of details.

[0130] For scenes containing dynamic elements, observe whether the movement trajectories of objects in the generated scene are natural and smooth and conform to the physical movement laws in the real scene, and quantitatively evaluate by calculating the smoothness of the position changes of objects in consecutive frames and the reasonableness of the acceleration.

[0131] Analyze whether the corresponding changes in lighting, object distribution, activity conditions, etc. are in line with the logic and common laws of the real scene when the scene changes from one state to another, and measure the degree of fit by setting some rules and counting the key feature changes during the state transition process.

[0132] If the scenario generation technology is applied to scenarios that support interaction, the timeliness and accuracy of the response of the generated scenario to the user's interactive operations and the rationality of the interactive effects are evaluated, and the analysis is performed by simulating user interactive operations and recording relevant response data.

[0133] For specific task scenarios, the success rate of users completing specific training tasks using the generated virtual scenarios is counted to measure whether the scenarios help users achieve their goals and whether they meet the actual needs of the application scenarios.

[0134] The validation set data is input into the overall model after quantum computing and QGAN adversarial training in sequence, allowing the model to generate corresponding scenario data according to the established process, and the generated data results are recorded to ensure that the data generation process is complete and free of abnormalities, providing accurate samples for subsequent evaluation and analysis.

[0135] Use the evaluation indicators determined previously to quantitatively evaluate the generated scene data and calculate the corresponding indicator values. For each generated scene image, calculate its PSNR and SSIM values, and count the values of indicators related to the coherence of object motion in the scene, etc., to form a complete set of indicator data, so as to intuitively understand the performance of the model in various aspects.

[0136] Compare and analyze the currently calculated indicator values with the model performance indicators before training (if recorded) or the results of the previous verification round, focusing on whether there is a significant performance improvement in visual quality, scene dynamic rules and other specific application scenarios after adding quantum computing and QGAN adversarial training, including observing whether the PSNR value is improved, whether the coherence of object motion is enhanced, etc., to find out the improvements of the model and possible deficiencies that may still exist through comparison.

[0137] If comparative analysis finds that the model performance has not improved significantly, and that various indicators are still far from the expected targets, it may be that the model has not fully converged due to insufficient training rounds. In this case, you can appropriately increase the training rounds, continue to train the model, and then perform verification and evaluation again to observe whether the performance has improved.

[0138] If it is found that the generated scenario data is abnormal in some aspects, the data interaction and parameter transfer methods between quantum computing and QGAN will be reviewed, and it will be checked whether the data format at each interface is correct and whether the parameter update is carried out according to the predetermined logic. Appropriate modifications will be made to the unreasonable places to ensure that the data can flow accurately and smoothly between different modules and ensure the collaborative working effect of the model.

[0139] When the analysis results indicate that there may be limitations in the overall architecture of the model, which affect the performance improvement, then try to adjust and optimize the network structure of the QGAN generator or discriminator by increasing or decreasing the number of layers, changing the neuron connection method, adjusting the convolutional kernel size, etc. After that, retrain and validate, and continuously iterate this process until the generator can produce new scene data that is realistic and conforms to the dynamic laws of the scene, meeting the specific application requirements.

[0140] Carefully study the data structures, function call methods, parameter passing requirements, etc. of the CNN-HMM hybrid model, QGAN model, and Babylon.js framework respectively. Clearly define the format, dimensions, etc. of the feature data output by the CNN-HMM hybrid model, the format of the scene data output by the QGAN generator, and at the same time understand the data types that Babylon.js framework can receive and the expected input interface forms, so as to determine the specific functions that the adaptation code needs to implement.

[0141] According to the above analysis, write code to perform data format conversion and mapping. If the features output by the CNN-HMM hybrid model are represented in a specific tensor format, while the data format required by the Babylon.js framework is a specific JSON format or other custom formats, functions need to be written to parse and convert the tensor data into the target format. For the scene data generated by QGAN, necessary format adjustments are also made to ensure that it can be correctly recognized and processed by the Babylon.js framework, ensuring that data can be smoothly transferred between the model and the framework without loss of key information.

[0142] Define the function call logic between the model and the Babylon.js framework, so that Babylon.js can call the model for scene generation or processing operations at the appropriate time. For example, when new scene elements need to be generated in the Babylon.js framework, write corresponding call functions to trigger the QGAN model to generate data; when the scene structure features need to be analyzed, call the CNN-HMM hybrid model to extract and return relevant feature information. Ensure that the parameter passing of these function calls is accurate and the return values can be correctly received and utilized by the framework.

[0143] In the Babylon.js framework, select some simple and representative scene elements for testing. For example, generate single basic geometric bodies or simple texture patterns and other scene elements, and use the embedded model to generate relevant data and transfer it to the framework for display processing.

[0144] Observe whether these simple elements generated by the observation model can be correctly displayed in the Babylon.js scene, and check whether their properties such as position, size, and orientation meet the expected settings, so as to preliminarily verify whether the functions of the model in generating and displaying basic elements are normal after being integrated with the framework.

[0145] In the direction of the framework passing parameters to the model, set different parameter values and check whether the model can correctly receive and perform corresponding processing and operations according to the parameter requirements. For example, change the random noise parameters passed to the QGAN model and observe whether the generated scene data will change reasonably accordingly; adjust the parameters related to the input image size passed to the CNN-HMM hybrid model and check whether the feature extraction results meet the expected corresponding changes.

[0146] In the direction of the model returning data to the framework, check whether the returned data can be smoothly received by the framework and correctly applied to the scene construction process. For example, verify whether the scene structure feature data returned by the CNN-HMM hybrid model can be used by the Babylon.js framework to reasonably layout the objects in the scene, and whether the scene data generated by QGAN can be accurately integrated into the overall scene rendering process to ensure the integrity and accuracy of data interaction.

[0147] Let the integrated system run continuously for a period of time to simulate long-term working scenarios that may occur in actual use, and observe whether there are stability problems such as system crashes, memory leaks, and software crashes. During this period, continuously perform scene generation, element addition, and interaction operations, etc., to increase the system load and test whether the system can run stably under different pressure conditions.

[0148] Artificially create some abnormal situations, such as inputting data that does not meet the requirements, interrupting the network connection (if network-related functions are involved), and suddenly changing the system resource configuration, etc., and check whether the system has corresponding exception handling mechanisms and can properly handle these situations without serious errors or data loss and other problems to ensure the stability and reliability of the entire integrated system.

[0149] Deeply analyze various scene elements generated by QGAN, including specific details such as the geometric shapes of objects, texture information, lighting effects, etc. Clearly identify the type of geometric body of the generated object model, the characteristics of its surface texture, and the built-in lighting rendering attributes. At the same time, refer to the analysis results of CNN-HMM on the scene structure and dynamic rules to determine the accurate positions, angles, and hierarchical relationships of these scene elements in the entire virtual scene. Based on the information about the relative positions of objects and scene layouts extracted by CNN-HMM, place the objects generated by QGAN at appropriate coordinate positions; adjust the rotation angles of the objects according to its analysis of object orientations and angle change trends to make them conform to the logic and visual rationality of the scene; then, according to the hierarchical structure characteristics of the scene, determine the hierarchical relationships of each element to ensure the layering and three-dimensional sense of the scene.

[0150] Select an appropriate resolution according to the target platform that the virtual scene ultimately faces and the specific application scenario. Consider the characteristics of different display devices, such as screen size and pixel density, to further optimize the resolution settings. For devices with a large screen size and high pixel density, appropriately increase the resolution to avoid image blurring or jaggedness; conversely, for devices with a small screen or limited performance, reasonably reduce the resolution to ensure smooth rendering and interactive response of the scene.

[0151] Determine an appropriate frame rate according to the dynamic complexity and interaction requirements of the scene. For scenes containing fast-moving objects and frequent interaction operations, a higher frame rate is required to ensure the smoothness of animations and interactions, and to avoid phenomena such as stuttering and motion blur, so that users have a smooth experience during operation and viewing; while for relatively static and less interactive scenes, appropriately reduce the frame rate to reduce the system rendering burden while ensuring the basic visual effects and improve the overall performance.

[0152] Comprehensively consider performance factors such as the graphics processing ability and memory of the hardware device during the process of optimizing the frame rate. Through performance testing and monitoring, observe the resource occupancy of the system under different frame rate settings, and find the best frame rate balance point that can meet the smoothness requirements and enable the device to operate stably, avoiding problems such as performance degradation or even crashes caused by excessive frame rates, or low frame rates affecting the user experience.

[0153] According to the characteristics of the scene and rendering requirements, select a suitable one or several anti-aliasing technologies (including super sampling anti-aliasing SSAA, multi-sampling anti-aliasing MSAA, fast approximate anti-aliasing FXAA, etc.) for use in combination. For scenes with rich details, a large number of fine lines and edges, you can use sampling-based anti-aliasing methods such as MSAA to reduce the jagged effect of the image edge by increasing the sampling points and improve the smoothness of the picture; and for scenes with high performance requirements and strong real-time performance, use FXAA, a relatively lightweight and fast anti-aliasing technology, to ensure a certain anti-aliasing effect without placing too much burden on rendering performance.

[0154] For the selected anti-aliasing technology, adjust the relevant parameters to optimize the anti-aliasing effect. During the parameter adjustment process, observe the changes in the scene image in real time, compare the jagged conditions of the image edges, the clarity and naturalness of the overall image under different parameter settings, etc., and find the parameter combination that can achieve the best anti-aliasing effect, so that the edges of objects in the scene look more natural and smooth, thereby improving the visual quality.

[0155] According to factors such as the scene's lighting conditions, object types, and rendering performance requirements, select the appropriate shadow generation algorithm (including shadow mapping, ray tracing shadows, etc.). For scenes with relatively simple lighting, a large number of objects, and high requirements for real-time rendering performance, the shadow mapping algorithm is used to simulate shadow effects by generating depth textures, and the calculation efficiency is relatively high; for high-end scenes that pursue extremely realistic light and shadow effects and hardware performance allows, ray tracing shadows are used, which can more accurately simulate light propagation and occlusion relationships and generate very realistic shadows, but the calculation cost is high and requires strong hardware support.

[0156] After selecting the shadow generation algorithm, further adjust the relevant parameters to optimize the quality and visual effects of the shadow. For example, improving the shadow resolution can make the shadow edge clearer and sharper, and reduce blur; properly adjusting the shadow offset can avoid the wrong rendering position of the shadow; adjusting the softness can make the shadow look more natural and soft, and more in line with the light and shadow effects in the real scene. Through these detailed optimizations, the shadows in the scene are more real and natural, and the realism and immersion of the entire scene are enhanced.

Claims

1. A method for generating a quantum-style AR scene based on a hybrid model under the Babylon.js framework, characterized in that, It includes the following steps: S1. Label the key object and scene category information for the collected scene data (image sequences and depth data containing rich visual and spatial information, ensuring coverage of different states and detailed changes of the scene), divide the training set, validation set, and test set proportionally, and make the data format meet the input requirements of the subsequent model; S2. Use VGGNet as the basic structure to construct a CNN, adjust relevant parameters according to the scene characteristics, connect the output features of the CNN to an HMM, regard the feature sequence extracted by the CNN as the observation value of the HMM, and set the hidden state of the HMM based on the internal logic of the scene; S3. Determine the initial transition probability and emission probability matrices, use the Baum-Welch algorithm to iteratively optimize in combination with the training set, learn the dynamic change rules of the scene, and at the same time input the training data in an end-to-end manner to co-train the CNN and the HMM. Monitor the performance indicators of the model with the validation set, and fine-tune the CNN structure and HMM parameters until the model performance meets the requirements; S4. Plan the number of qubits and the quantum gate sequence as needed to construct a quantum circuit, input it into the Qiskit Aer quantum simulator, simulate the state evolution of the qubits and quantum gate operations, and observe the impact of the simulation results on the generated scene data and perform preliminary optimization on it; S5. Classify the scene implicit features extracted by the CNN-HMM hybrid model. Use a classical computer to complete simple feature extraction and preliminary processing, input the preprocessed features into the quantum circuit, and use quantum computing for more complex feature fusion or generate new features with quantum characteristics; S6. Define the QGAN generator and discriminator network architectures, use a small-scale dataset to verify and integrate the results obtained by quantum computing and the classical computing part, use the integrated information as part of the input of the QGAN generator, and generate "fake" scene data in combination with random noise; S7. The discriminator discriminates between real and generated scene data, and the two adversarial games optimize their respective parameters. During the training stage, feed the real scene samples and the generator output to the discriminator alternately every 5-10 rounds, and update the weights of the generator and discriminator according to the discrimination results through backpropagation; S8. After completing the training, use the validation set again to verify the improvement of the overall model performance after adding quantum computing and QGAN adversarial training. After multiple rounds of iteration, prompt the generator to produce new scene data that is realistic and conforms to the dynamic rules of the scene; S9. Embed the trained CNN-HMM hybrid model and QGAN model into the Babylon.js environment, ensure smooth data interaction according to the adaptation code, enable Babylon.js to call the generation or processing results of the model to participate in the scene construction process, and generate tests to check the functions after the integration of the model and the framework; S10. Within the framework, combine the scene elements generated by the QGAN with the CNN-HMM analysis results, and use the Babylon.js framework to completely build the virtual scene, accurately set the object coordinates, angles, and hierarchical relationships. For display requirements, optimize the rendering parameters, frame rate, and resolution, enable anti-aliasing and shadow optimization technologies to optimize the picture and light and shadow visual effects, and restore the real scene and reflect the predicted changes.

2. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that, The annotation and division of the scene data described in step S1 include: Collect scene image sequences and depth data under different environmental conditions, time points, and scene layouts. The data covers various states of the scene, including the scene under different light intensities, different placement positions and angles of objects, dynamic changes of the objects in the scene, and different spatial scale information, ensuring that the diversity and complexity of the real scene can be comprehensively reflected; Annotate the collected data. For the key objects in the image sequence, use the bounding box or pixel-level annotation method to accurately mark the position and contour of each key object and assign its corresponding category label. At the same time, also perform category annotation on the entire scene; Randomly divide the annotated data set into a training set, a validation set, and a test set according to a certain ratio. During the division process, ensure that the data distributions in different sets are as similar as possible to avoid the adverse effects of data deviation on model training and evaluation. At the same time, perform format conversion and preprocessing on the data to make it meet the input requirements of subsequent models such as CNN and HMM.

3. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that The construction of the CNN and the integration and initialization of the HMM under the VGGNet structure described in step S2 include: Build the CNN based on the VGGNet as the basic structure. First, determine the number of layers of the network and appropriately adjust the number of layers of the classic structure of the VGGNet according to the complexity of the scene data; Determine the ReLU function as the activation function to effectively solve the problem of gradient disappearance and speed up the training speed of the network. At the same time, configure the pooling layer to reduce the resolution of the feature map, reduce the amount of data, and extract the main features. In the last few layers of the network, set the fully connected layer according to the task requirements. The number of neurons in the fully connected layer can be adjusted according to the number of scene categories or feature dimensions, and is used to map the features extracted by the convolutional layer to the final output category or feature representation; Connect the features output by the CNN to the HMM, determine the observations of the HMM, that is, use the feature sequence extracted by the CNN as the observation sequence of the HMM, and set the hidden state of the HMM based on factors such as the internal logic of the scene, the state changes of the objects in the scene, and the transitions of the scene. The hidden state reflects the dynamic change rules behind the scene and is associated with the feature sequence extracted by the CNN, laying a foundation for subsequent learning of the scene dynamic change rules.

4. The quantum-style AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that The CNN-HMM collaborative training described in step S3 includes: Determine the initial transition probability matrix and emission probability matrix of the HMM. The transition probability matrix represents the transition probabilities between hidden states, and the emission probability matrix represents the probabilities of observing specific features in each hidden state. Initially, set the probability values according to prior knowledge or a simple uniform distribution assumption, and these initial values will be continuously updated in subsequent iterative optimizations; Use the Baum-Welch algorithm to iteratively optimize the parameters of the HMM in combination with the training set. In each iteration, first calculate the probability distribution of the observation sequence in each hidden state according to the current model parameters (the transition probability matrix and the emission probability matrix) (E-step), and then re-estimate the model parameters according to the probability distribution to maximize the likelihood function (M-step). During this process, input the training data in an end-to-end manner to co-train the CNN and the HMM; When training the CNN, use the feedback information of the HMM as an additional supervision signal to adjust the weights of the CNN, so that the features extracted by the CNN are more conducive to the HMM to model the dynamic changes of the scenario. Conversely, in the iterative optimization process of the HMM, use the more accurate feature sequences extracted by the CNN to continuously update the transition probability matrix and the emission probability matrix; After completing a certain number of rounds of the training, use the validation set to verify the model, calculate the performance metrics of the model on the validation set, and according to the verification results, finely tune the structural parameters of the CNN and the parameters of the HMM, and repeat the training and verification process until the model performance meets the initial requirements.

5. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, wherein The quantum circuit construction and simulation optimization in step S4 includes: Determine the number of qubits according to the complexity of the scenario data and the dimensions and characteristics of the features extracted by the CNN-HMM hybrid model. When making a preliminary estimate, start with a smaller number of qubits and then gradually increase the number of qubits according to subsequent simulation results and the evaluation of the accuracy of the feature representation; Construct a quantum gate sequence based on the basic gate operations of quantum computing to achieve specific transformations and processing of the input features, achieving the effect of complex feature fusion or generating new features. First, use single-bit gates to initialize the qubits and set them to specific initial states, and then use multi-bit gates to achieve entanglement and interaction between qubits, thereby simulating the complex correlations between the features; Input the constructed quantum circuit into the Qiskit Aer quantum simulator. In the simulator, set the simulation parameters. After starting the simulation, observe the state evolution process of the qubits, that is, the process in which the qubits gradually change from the initial state to the final state as the quantum gate operations are sequentially executed, and analyze the impact of the quantum gate operations on the qubit states; According to the simulation results, adjust the parameters and structure of the quantum circuit, change parameters such as the operation angles of the quantum gates or the action order of the gates, attempt to add or delete certain quantum gates, or change the connection manner between the quantum gates to adjust the structure. Through continuous adjustment and re-simulation, optimize the quantum circuit so that it can better process the scene features and improve the ability to generate new features with quantum characteristics and meeting the scene requirements.

6. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, wherein The feature classification and quantum processing in step S5 include: Classify the scene implicit features extracted by the CNN-HMM hybrid model according to the type of the features, the importance of the features, or the correlation between the features and different aspects of the scene. Classify some simple and low-level features into one category, and these features can be preliminarily processed through simple mathematical operations on a classical computer. And classify some features related to the overall structure of the scene or the complex relationships between objects into another category, and these features will be input into the quantum circuit for more complex processing; For simple feature extraction and preliminary processing, operate on a classical computer, then encode and convert the features after classical preprocessing according to the input requirements of the quantum circuit, and convert the classical data into a representation form of a quantum state. Map the feature vector to the state of the quantum bit through a specific encoding algorithm to ensure that the quantum circuit can correctly receive and process this feature information; Input the converted features into the quantum circuit for processing. The quantum circuit utilizes the superposition state and entanglement characteristics of the quantum bits, and performs complex feature fusion operations on the input features through the quantum gates, enabling entanglement to occur between the features encoded by different quantum bits, thereby simulating complex feature interactions that are difficult to achieve in classical computing.

7. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, wherein The QGAN architecture determination and data integration in step S6 include: Use a small-scale dataset to verify and integrate the results obtained from quantum computing and the results of the classical computing part. First, run the codes of the quantum computing and classical computing parts respectively on the small-scale dataset to obtain their respective output results. Then, evaluate the quality and effectiveness of the results through some verification metrics; According to the verification results, integrate the results of the quantum computing and classical computing. If it is found that some of the quantum computing results are insufficient in certain aspects, make appropriate adjustments or corrections to them, and use the integrated information as part of the input of the QGAN generator to generate "fake" scene data in combination with random noise. The addition of the random noise can increase the diversity and randomness of the generated data, enabling the generator to explore a wider scene data space and avoiding generating overly single or deterministic scene data; During the generation process, ensure that the integrated input information and the random noise can be reasonably processed and fused in the network of the generator to generate high-quality "fake" scene data, laying a foundation for subsequent QGAN training and scene construction.

8. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that The QGAN adversarial training described in step S7 includes: Before starting the training, first initialize the network parameters of the generator and discriminator of the QGAN. The generator parameters are randomly initialized with relatively small random values to avoid overly large gradient updates at the beginning of the training. The discriminator parameters are initialized similarly. At the same time, prepare a representative dataset of real-scenario samples, which should cover various scenario states and types and match the output of the generator in terms of data format and feature representation for effective comparison and discrimination. Enter the training loop. Data feeding and parameter updates are performed every 5 - 10 rounds. In each round, first, the generator generates a batch of "fake" scenario data based on its current parameters and input. Then, the real-scenario samples and the generator output are alternately fed to the discriminator, which discriminates these data and outputs the probability estimates of whether each data sample is real-scenario data or generated data. Calculate the loss function value of the discriminator according to the discrimination results. The binary cross-entropy loss function is used. For real samples, the discriminator outputs a probability close to 1, and for generated samples, it outputs a probability close to 0. Calculate the gradient of the discriminator loss function with respect to its parameters through the backpropagation algorithm and use the Adam optimizer to update the discriminator parameters so that the discriminator can better distinguish real and generated data. Next, use the generator to generate another batch of "fake" scenario data and input it into the discriminator with the updated parameters. This time, calculate the loss function of the generator to make the discrimination result of the discriminator for the generated data close to 1, that is, the generated data can "fool" the discriminator. Similarly, calculate the gradient of the generator loss function with respect to its parameters through backpropagation and update the generator parameters. In this way, the generator and the discriminator continuously optimize their own parameters during the adversarial process and gradually improve the ability of the generator to generate realistic scenario data.

9. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that, The model performance verification and iterative optimization described in step S8 include: After completing a certain number of rounds of adversarial training, prepare a validation set. The validation set is independent and similar to the training set and the test set in terms of data distribution and contains samples of different scenario categories and states. Determine the metrics for evaluating the model performance, including visual quality metrics for scenario generation, the degree of fit of scenario dynamic laws, and other metrics related to specific application scenarios. Input the validation set data into the overall model after the quantum computing and QGAN adversarial training to obtain the generated scenario data. Evaluate the generated data using the pre-set metrics, calculate the values of each metric and compare them with the results before training or in the previous validation rounds to analyze the improvement in the performance of the overall model after adding the quantum computing and QGAN adversarial training. According to the verification results, adjust and optimize the model. If the number of training rounds is insufficient, increase the number of training rounds; if it is an integration problem, re-examine the data interaction and parameter transfer methods between the quantum computing and the QGAN, and make appropriate modifications; if it is a model architecture problem, try to adjust the network structure of the QGAN generator or the discriminator by increasing or decreasing the number of layers, changing the neuron connection method, etc., and then perform training and verification again. After multiple rounds of iteration, until the generator can produce new scene data that is realistic and conforms to the dynamic laws of the scene, meeting the application requirements.

10. The quantum AR scene generation method based on the hybrid model under the Babylon.js framework according to claim 1, characterized in that The model embedding and integration testing in step S9 includes: Integrate the trained CNN-HMM hybrid model and the QGAN model into the Babylon.js environment. First, write adaptation code to ensure that the code structure and data interface of the model are compatible with the Babylon.js framework. At the same time, handle the function call relationship between the model and Babylon.js so that Babylon.js can call the model at the appropriate time for scene generation or processing operations; Conduct some simple scene element generation tests to check the functions after the model is integrated with the framework. Observe whether the elements generated by the model can be correctly displayed in the Babylon.js scene, check whether their attributes such as position, size, and direction meet the expectations, verify whether the data interaction between the model and the framework is smooth, and whether the parameters passed from Babylon.js to the model can be correctly received and processed by the model. At the same time, check whether the model can work properly in the new Babylon.js environment to ensure the stability and reliability of the entire integrated system.

Citation Information

Cited By

  • Seal removing and document repairing method based on quantum state cooperative regulation and control

    CN120931533A

  • Method for removing seal and repairing document based on quantum state cooperative regulation

    CN120931533B