Perception enhancement method and system based on environment, electronic equipment and storage medium
Through the deep convolutional neural network and digital twin model of AR devices, combined with ray tracing algorithms and knowledge graphs, intelligent display parameters adjustment of AR devices in complex environments is realized, solving the problem of insufficient display adaptability and improving display stability and user experience.
Patent Information
- Application Number
- CN202510456794.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-12
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AR devices have insufficient display adaptability in complex environments and cannot comprehensively consider a variety of environmental factors, resulting in a decrease in display effect and it is difficult to optimize targetedly according to the characteristics of specific scenarios.
The environment image is obtained through the AR device camera, the scene classification is performed using a deep convolutional neural network, the scene type transfer probability matrix and Markov prediction model are combined for dynamic optimization, the knowledge graph is constructed for template matching, and the environment parameter regulation is used for ray tracing algorithm and digital twin model to realize intelligent adjustment of AR display parameters.
It improves the display stability and user experience of AR devices in complex environments, realizes accurate understanding of the scene and precise regulation of environmental parameters, and provides timely risk warnings and operational response suggestions.
Smart Images

Figure CN120298633A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of AR, and particularly to a method, system, electronic device, and storage medium for enhanced perception based on the environment. Background Art
[0002] Augmented Reality (AR) technology enhances the user's perceptual experience by superimposing virtual information on real-world scenes and has been widely applied in fields such as industrial manufacturing, medical surgery, and education and training. In AR applications, to ensure the clear visibility of virtual information, AR devices need to adjust display parameters according to different environmental conditions. Currently, mainstream AR devices usually adopt a method of adaptively adjusting display parameters based on an ambient light sensor, that is, dynamically adjusting the display brightness and contrast by measuring the ambient light intensity. However, this method only considers a single factor of ambient light and ignores the influence of other environmental parameters on the AR display effect. For example, in a high-temperature or high-humidity environment, the performance of the optical system of the AR device may change, resulting in a decline in the display effect; in a complex scene, environmental occluders will affect the visibility of virtual content, and it is difficult to accurately perceive these factors solely relying on the light sensor. In addition, the requirements for the AR display effect also vary in different application scenarios, and it is difficult for the existing technology to perform targeted display optimization according to the characteristics of specific scenarios. Therefore, it is necessary to develop a technical solution that can comprehensively consider multiple environmental factors and intelligently adjust AR display parameters according to the characteristics of the scenario. Summary of the Invention
[0003] This application provides a method, system, electronic device, and storage medium for enhanced perception based on the environment, which is used to intelligently adjust AR display parameters according to the characteristics of the scenario and solve the problem that it is difficult for the existing technology to perform targeted display optimization according to the characteristics of specific scenarios.
[0004] In a first aspect, this application provides a method for enhanced perception based on the environment, including: Obtaining N environmental images through the camera of the AR device to obtain an environmental image set, and performing scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type; Performing optimization processing on the preliminary scene type to obtain a target scene type, and matching a corresponding application scene template from a preset scene template library according to the scene type; Determining the type of environmental parameters to be collected according to the application scene template, and collecting environmental data through the sensor corresponding to the type of environmental parameters to be collected to obtain target environmental data; Generating AR device adjustment parameters according to the target environmental data; Adjusting the display parameters of the AR device according to the AR device adjustment parameters.
[0005] In the above technical solution, by constructing a closed-loop system for intelligent adjustment of AR display parameters from environmental perception, the key technical problem of insufficient display adaptability of traditional AR devices in complex environments is significantly solved. First, the system obtains an environmental image set through the AR device camera, and realizes intelligent classification of scenes based on a deep convolutional neural network, overcoming the limitation of traditional image recognition methods in insufficient understanding of complex scenes. By introducing a scene type transition probability matrix and a Markov prediction model, the system realizes dynamic optimization of the scene types of continuous image frames, significantly improving the continuity and accuracy of scene understanding compared with traditional single-frame scene recognition methods.
[0006] The semantic matching of scene templates is the core innovation of the system. By constructing a knowledge graph containing scene types, scene attributes, and scene parameters, and using the cosine similarity and softmax normalization methods, the system can accurately match the most suitable application scene template. This semantic-based template matching method breaks through the limitations of traditional empirical scene adaptation and realizes the intelligence and accuracy of scene template selection. Based on the matched application scene template, the system dynamically determines the types of key environmental parameters, and through multi-sensor collaborative acquisition, a comprehensive and accurate environmental data acquisition model is established.
[0007] The construction of the digital twin model is the key technology for the system to realize intelligent regulation of environmental parameters. Through the joint representation of the three-dimensional geometric model and the physical property model, the system not only visualizes the environment but also can perform complex physical property simulations. The ray tracing algorithm accurately calculates the environmental light occlusion factor, providing an accurate physical basis for the dynamic adjustment of the AR content transparency parameter. The calculation of the transparency parameter τ takes into account the environmental light intensity and the maximum brightness of the device to ensure the best visibility of virtual information under different lighting conditions.
[0008] The influence of environmental parameters on display performance is a key factor ignored by traditional AR devices. In this application, by analyzing the microscopic influence of physical parameters such as temperature and humidity on the optical system, a mathematical model of the degradation of display performance with environmental parameters is established, and the contrast adjustment coefficient is calculated accordingly. This display parameter compensation method based on a physical model significantly improves the display stability and user experience of AR devices in complex environments.
[0009] The risk assessment and early warning mechanism is another important innovation of the system. By constructing a multi-dimensional and multi-level risk determination algorithm, the system can accurately evaluate the risk level of environmental parameters and generate warning information with precise semantics and easy to understand. This intelligent risk perception method not only provides timely environmental risk warnings but also provides users with actionable risk response suggestions.
[0010] In the second aspect of this application, a perception enhancement system based on the environment is provided, and the system includes: A data acquisition module, configured to obtain N environmental images through a camera of an AR device, obtain an environmental image set, and perform scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type; A template matching module, configured to perform optimization processing on the preliminary scene type to obtain a target scene type, and match a corresponding application scene template from a preset scene template library according to the scene type; An environmental parameter determination module, configured to determine the type of environmental parameters to be collected according to the application scene template, and collect environmental data through a sensor corresponding to the type of environmental parameters to be collected to obtain target environmental data; An adjustment parameter determination module, configured to generate AR device adjustment parameters according to the target environmental data; An adjustment module, configured to adjust the display parameters of the AR device according to the AR device adjustment parameters.
[0011] In a third aspect of the present application, a computer storage medium is provided. The computer storage medium stores multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the above method steps.
[0012] In a fourth aspect of the present application, an electronic device is provided, including a processor, a memory, a user interface, and a network interface. The memory is used to store instructions, the user interface and the network interface are used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device executes the above method.
[0013] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. By constructing a closed-loop system for intelligent adjustment of AR display parameters from environmental perception, the present application significantly solves the key technical problem of insufficient display adaptability of traditional AR devices in complex environments. The system first obtains an environmental image set through the AR device camera and realizes intelligent classification of scenes based on a deep convolutional neural network, overcoming the limitation of traditional image recognition methods in insufficient understanding of complex scenes. By introducing a scene type transition probability matrix and a Markov prediction model, the system realizes dynamic optimization of the scene types of continuous image frames, and significantly improves the continuity and accuracy of scene understanding compared with traditional single-frame scene recognition methods.
[0014] 2. By constructing a knowledge graph including scene types, scene attributes, and scene parameters, and adopting the cosine similarity and softmax normalization methods, the system can accurately match the most suitable application scene template. This semantic-based template matching method breaks through the limitation of traditional empirical scene adaptation and realizes the intelligence and accuracy of scene template selection. Based on the matched application scene template, the system dynamically determines the types of key environmental parameters and establishes a comprehensive and accurate environmental data acquisition model through multi-sensor collaborative acquisition. Brief Description of the Drawings
[0015] Figure 1 It is a schematic flowchart of a method for enhancing perception based on the environment provided by an embodiment of the present application; Figure 2 It is an architecture diagram of a system for enhancing perception based on the environment provided by an embodiment of the present application; Figure 3 It is a schematic structural diagram of an electronic device provided by the present application. Detailed Embodiments
[0016] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0017] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, using words such as "for example" or "for instance" aims to present relevant concepts in a specific manner.
[0018] In the description of the embodiments of the present application, the meaning of the term "plurality" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0019] The technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0020] Based on the above background technology, further, please refer to Figure 1 , Figure 1The figure is a schematic flowchart of a method for enhancing perception based on the environment provided by an embodiment of the present application. This system can be implemented relying on a computer program or run as an independent tool-like application. Specifically, in the embodiment of the present application, this method can be applied to a server, but can also be applied to electronic devices such as a server. A method for enhancing perception based on the environment includes the following steps: S101, obtain N environmental images through the camera of the AR device to obtain an environmental image set, and perform scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type; Specifically, since environmental images may vary at different times, in order to accurately identify the scene type, it is necessary to obtain multiple environmental images through the camera of the AR device. Specifically, the camera of the AR device obtains an environmental image at every preset time interval, and continuously obtains N environmental images, where N is a preset positive integer, to obtain an environmental image set. When performing scene classification based on the environmental images in the environmental image set, first perform size normalization processing on the environmental images in the environmental image set to uniformly adjust environmental images of different sizes into standardized images of a preset size, and perform illumination compensation and noise filtering on the standardized images to obtain enhanced images. Then use a pre-trained deep convolutional neural network to extract multi-scale feature maps of the enhanced images. The deep convolutional neural network extracts the feature information of the images layer by layer through multiple convolutional and pooling operations. Stack spatial attention weights on the multi-scale feature maps to highlight the feature expressions of important regions and obtain enhanced feature maps. Then perform feature dimensionality reduction on the enhanced feature maps, compress the high-dimensional features into low-dimensional feature vectors, and input the feature vectors into a preset classifier to obtain the prediction probabilities of each scene type. Finally, perform softmax normalization processing on the prediction probabilities to convert the prediction probabilities into a probability distribution, and screen the probability distribution according to a preset confidence threshold to screen out the scene type with the highest confidence as the preliminary scene type. Through the above steps, it is possible to effectively remove the noise interference in the environmental images, extract the discriminative features of the images, and improve the accuracy of scene classification.
[0021] Based on the above embodiment, as an alternative embodiment, performing scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type includes: S201, perform size normalization processing on the environmental images in the environmental image set to obtain standardized images of a preset size, and perform illumination compensation and noise filtering on the standardized images to obtain enhanced images; Specifically, when performing scene classification on environmental images, it is first necessary to preprocess the environmental images in the environmental image set. Specifically, size normalization processing needs to be performed to uniformly scale the original environmental images of different sizes to a preset standard size to obtain normalized images. The reason for performing size normalization is that subsequent feature extraction and classification algorithms usually require input images to have the same size. The standardized size can be set according to specific application requirements and model complexity. Common ones are 224x224, 256x256, 299x299, etc.
[0022] After obtaining the normalized image, it is also necessary to perform image enhancement on it. Here, two enhancement methods of light compensation and noise filtering are adopted. Light compensation is to reduce the impact of environmental light changes on image quality. Since the original environmental images may be taken under different lighting conditions and there are differences in brightness and contrast, directly using them for classification may affect the accuracy. Common light compensation methods include histogram equalization, Retinex algorithm, etc., which can correct the uneven lighting problem to a certain extent and make the image clearer and more uniform.
[0023] Noise filtering is to remove the noise interference in the image and improve the image quality. Due to the influence of the imaging device and shooting environment, environmental images often have certain noise, such as Gaussian noise, salt-and-pepper noise, etc., which will affect feature extraction and classification performance. Common noise filtering methods include median filtering, bilateral filtering, wavelet transform, etc., which can effectively suppress noise while retaining the image edge and texture details.
[0024] After light compensation and noise filtering, the enhanced environmental image is obtained. Compared with the original normalized image, it has higher image quality and visual effect, which is more conducive to subsequent feature extraction and scene classification tasks. Such preprocessing can significantly improve the robustness and generalization ability of the classification algorithm and achieve better recognition effects in complex and changing real environments.
[0025] S202, use the pre-trained deep convolutional neural network to extract the multi-scale feature maps of the enhanced image, and superimpose the spatial attention weights on the multi-scale feature maps to obtain the enhanced feature maps; Specifically, after obtaining the preprocessed enhanced image, the next step is to extract the feature representation of the image to prepare for subsequent scene classification. Here, a pre-trained deep convolutional neural network is used to complete the feature extraction task.
[0026] Deep convolutional neural network is a powerful deep learning model. By stacking operations such as convolution, pooling, and activation layer by layer, it can automatically learn and extract hierarchical features in images. Different from traditional manually designed features, deep convolutional neural network can learn more abstract and semantic feature representations end-to-end, which has significant advantages for the recognition of complex scenes.
[0027] The deep convolutional neural network used here is pre-trained, that is, the model weights obtained by pre-training on a large-scale image dataset. Common pre-trained models include AlexNet, VGGNet, ResNet, etc. The advantage of using a pre-trained model is that it can directly utilize the general features it has learned without having to train from scratch, which greatly saves time and computing resources. By fine-tuning on the target task, the pre-trained model can be well transferred to new application scenarios.
[0028] When using the pre-trained deep convolutional neural network to extract features, this application not only focuses on the features of the output layer of the network, but also extracts the feature maps of the intermediate layers. This is because the convolutional neural network has hierarchical feature learning ability. Shallow feature maps capture local and low-level features such as edges and textures; deep feature maps capture global and high-level semantic features. By fusing feature maps at different levels, a more comprehensive and robust image representation can be obtained. This extraction strategy of multi-scale feature maps has achieved good results in many visual tasks.
[0029] However, the feature maps learned by the convolutional neural network usually only focus on the saliency of local regions and ignore the global relationships between different regions. To further enhance the representation ability of the feature maps, a spatial attention mechanism is introduced here. Spatial attention can adaptively adjust the spatial distribution of the feature maps by learning the importance weights of different positions in the feature maps, highlighting the contributions of key regions and suppressing the interference of background noise.
[0030] S203, perform feature dimensionality reduction on the enhanced feature map to obtain a feature vector, and input the feature vector into a preset classifier to obtain the prediction probabilities of each scene type; Specifically, after obtaining the enhanced feature map, in order to further improve the classification efficiency and generalization performance, this application needs to perform dimensionality reduction processing on the feature map. The reason for performing feature dimensionality reduction is that the enhanced feature map usually has a high dimension, and directly using it for classification may bring problems such as large computational overhead and high overfitting risk. By dimensionality reduction, this application can significantly reduce the redundancy and noise of the feature representation while retaining the main feature information.
[0031] Common feature dimensionality reduction methods include principal component analysis (PCA), linear discriminant analysis (LDA), etc. Here, PCA is used for dimensionality reduction in this application. PCA maps the original high-dimensional features to a low-dimensional subspace through a linear transformation, maximizing the projection variance of the transformed features on the subspace. This can ensure that the features after dimensionality reduction retain the most discriminative information in the original data.
[0032] Specifically, in this application, the enhanced feature map is first flattened into a high-dimensional vector, and then a covariance matrix is constructed for the feature vectors of all samples. By performing eigenvalue decomposition on the covariance matrix, this application can obtain the eigenvectors and corresponding eigenvalues of the principal components. According to the magnitudes of the eigenvalues, the top k largest eigenvectors are selected as the principal components to form a dimensionality reduction matrix. Finally, multiplying the original feature vectors by the dimensionality reduction matrix can obtain the low-dimensional feature vectors after dimensionality reduction.
[0033] The feature vectors after dimensionality reduction not only have greatly reduced dimensions but also retain the most discriminative information in the original features. This can significantly improve the training and prediction efficiency of subsequent classifiers, while also reducing the risk of overfitting and improving the generalization ability of the model.
[0034] After obtaining the feature vectors after dimensionality reduction, this application inputs them into a preset classifier for scene classification. Common classifiers include support vector machine (SVM), random forest, logistic regression, etc. Here, SVM is selected as the classifier in this application. SVM separates samples of different classes by finding a maximum margin hyperplane in the feature space and has good classification performance and generalization ability.
[0035] This application inputs the feature vectors after dimensionality reduction into the trained SVM classifier. By calculating the distance from the sample to the hyperplane, the prediction probabilities of the sample belonging to each scene class can be obtained. These prediction probabilities reflect the similarity degree of the sample to each class, and the larger the probability value, the more likely the sample belongs to that class.
[0036] Through the above steps of feature dimensionality reduction and classification prediction, this application finally obtains the prediction probabilities of each environmental image belonging to various scene types. This method based on deep features and machine learning classifiers can effectively mine the scene semantic information in the image and achieve accurate recognition of complex environments. Compared with traditional handcrafted features and shallow classifiers, this method can often achieve better performance in scene classification tasks and has stronger environmental adaptability and robustness.
[0037] S204, perform softmax normalization on the prediction probabilities to obtain a probability distribution, and screen the probability distribution according to a preset confidence threshold to obtain a preliminary scene type.
[0038] Specifically, after obtaining the predicted probabilities of each scene type, the present application needs to post-process the probability values to obtain the final scene classification result. First, the present application performs softmax normalization on the predicted probabilities. Softmax is a commonly used normalization function that can transform a set of real numbers into a probability distribution. It scales the probability values to the range [0, 1] by taking the exponential of the predicted probabilities and dividing by the sum of the exponentials, and ensures that the sum of the probabilities of all classes is 1.
[0039] This normalization process has several benefits: First, it transforms the prediction results into a reasonable probability distribution, making the probability values between different classes comparable; Second, it amplifies the probability differences between different classes, helping to highlight the truly dominant classes; Finally, the normalized probability values are more stable and robust and are not easily affected by the numerical scale.
[0040] After obtaining the normalized probability distribution, the present application needs to screen the probability distribution according to a preset confidence threshold. The confidence threshold is a manually set parameter used to judge the credibility of the classification result. Specifically, the present application takes the classes in the probability distribution that exceed the confidence threshold as candidate types, and filters out the remaining classes. This can effectively remove some noise classes with low confidence and improve the accuracy and reliability of scene classification.
[0041] For example, if the present application sets the confidence threshold to 0.6, then only the classes with probability values greater than 0.6 will be regarded as valid scene types. If the predicted probability distribution of an image is {road: 0.8, building: 0.15, vegetation: 0.05}, then when the threshold is 0.6, only the "road" class will be retained, and the other two classes will be filtered out.
[0042] Through the screening of the confidence threshold, the present application obtains the preliminary scene type. This preliminary scene type reflects the scene semantics that the image is most likely to belong to, and has high confidence and reliability. In subsequent processing, the present application can also further optimize and refine the preliminary scene type using temporal information, prior knowledge, etc. to obtain the final scene classification result.
[0043] S102. Optimize the preliminary scene type to obtain the target scene type, and match the corresponding application scene template from the preset scene template library according to the scene type; Specifically, after obtaining the preliminary scene type, in order to further improve the accuracy and reliability of scene recognition, this application needs to optimize the preliminary scene type to obtain the final target scene type. The reason for scene type optimization is that the preliminary scene type is predicted based on a single-frame image, which may have certain misjudgments and instabilities. By introducing temporal information and prior knowledge, this application can smooth and correct the scene type on a longer time scale to obtain more accurate and coherent results.
[0044] Specifically, scene type optimization can be divided into the following steps: First, based on the preliminary scene type, this application analyzes a continuous image sequence within a period of time (such as several seconds). By counting the preliminary scene types of each frame of the image, this application can obtain a temporal distribution of the scene types. Then, this application uses temporal smoothing algorithms, such as Kalman filtering, hidden Markov models, etc., to smooth and denoise the temporal distribution. This can eliminate the random fluctuations of single-frame prediction and obtain a more stable and coherent sequence of scene types.
[0045] Next, this application can further optimize the scene type using prior knowledge. Prior knowledge includes probability models of scene transitions, spatio-temporal relationships of different scenes, etc. For example, this application can construct a scene transition probability matrix to characterize the prior probabilities of transitions between different scene types. Using this probability matrix, this application can perform maximum a posteriori probability estimation on the sequence of scene types after temporal smoothing to obtain the optimal sequence of scene types. This can make full use of the correlation information between scenes and improve the accuracy of scene recognition.
[0046] After temporal smoothing and optimization using prior knowledge, this application finally obtains the target scene type. Compared with the preliminary scene type, the target scene type has higher accuracy, coherence, and robustness, and can better reflect the true semantics of the environment. This provides reliable scene context information for subsequent AR content presentation and interaction.
[0047] After obtaining the target scene type, this application can match the corresponding application scene template from the preset scene template library. The scene template library is a pre-constructed knowledge base that contains information such as feature descriptions, interaction logics, UI layouts, etc. of various typical scenes. Each scene template corresponds to a specific AR application, such as indoor navigation, outdoor games, industrial maintenance, etc.
[0048] To achieve fast matching of scene templates, this application can organize the scene template library into an efficient data structure, such as a hash table, KD tree, etc. Using the target scene type as an index, this application can quickly retrieve the corresponding application scene template. At the same time, this application can also utilize the semantic relationships between scenes to construct a scene ontology to achieve semantic-based fuzzy matching. In this way, even if the user's environment has certain differences from the typical scene, similar application scene templates can still be matched.
[0049] Based on the above embodiments, as an alternative embodiment, the preliminary scene type is optimized to obtain the target scene type, including: S301, determining the preliminary scene type of N consecutive frames of environmental images in the environmental image set based on the preliminary scene type, where N is a preset positive integer greater than 1; Specifically, when optimizing the preliminary scene type, this application first needs to determine the preliminary scene type of N consecutive frames of environmental images in the environmental image set based on the preliminary scene type. Here, N is a preset positive integer greater than 1, representing the length of the time window considered in this application. The reason for selecting N consecutive frames of images is that the scene type of a single image may be misjudged and unstable, while multiple consecutive frames of images can provide more temporal context information, which helps to optimize and smooth the scene type.
[0050] Specifically, this application can adopt a sliding window method. Taking the current image as the center, (N - 1) / 2 frames of images are taken forward and backward respectively to form an image sequence with a length of N. Then, for each frame of image in this image sequence, it is labeled with the preliminary scene type obtained previously. In this way, this application obtains a time series of scene types, reflecting the changes in the scene type in a short period of time.
[0051] For example, assume that this application sets N to 5 and the preliminary scene type of the current image is "indoor". Then this application takes the first 2 frames and the last 2 frames of the current image, together with the current image, to form a 5-frame image sequence. Each of these 5 frames of images is labeled with the scene type, and a time series similar to ["outdoor", "indoor", "indoor", "indoor", "outdoor"] may be obtained. This time series indicates that there may be some fluctuations and uncertainties in the scene type in the short-time neighborhood of the current image.
[0052] By obtaining the preliminary scene types of consecutive N frames of environmental images, the present application actually transforms the prediction problem of scene types into a time series analysis problem. This provides important input information for subsequent optimization processing, enabling the present application to smooth and correct the scene types in the time dimension, eliminate the random noise of single-frame prediction, and obtain more stable and coherent results. At the same time, reasonably selecting the time window length N can capture the dynamic characteristics of scene changes without introducing excessive computational overhead, which is a key factor in scene type optimization.
[0053] S302. Calculate the frequency of occurrence of each scene type in the N frames of environmental images and normalize the frequency to obtain the scene type probability. Specifically, after determining the preliminary scene types of consecutive N frames of environmental images, the present application needs to further calculate the frequency of occurrence of each scene type and normalize the frequency to obtain the scene type probability. The reason for this step is that the frequency distribution of scene types can reflect the dominant degree of different scenes in the time series and provide important prior information for subsequent scene type optimization.
[0054] Specifically, the present application first traverses the preliminary scene type labels of these N frames of environmental images, counts the number of occurrences of each scene type, and obtains a frequency distribution. For example, for an N-frame sequence containing three scene types: "indoor", "outdoor", and "traffic", the present application may obtain such a frequency distribution: {"indoor": 3, "outdoor": 1, "traffic": 1}. This distribution indicates that within the current time window, the "indoor" scene appears the most times and is dominant.
[0055] However, the frequency distribution itself cannot be directly used for probability calculation and inference because different N values will cause changes in the frequency values. To eliminate this influence, the present application needs to normalize the frequency distribution. The purpose of normalization is to transform the frequency into a probability distribution such that the sum of the probabilities of all scene types is 1.
[0056] The normalization process can be achieved by dividing the frequency of each scene type by N. In the above example, the normalized scene type probability distribution is: {"indoor": 0.6, "outdoor": 0.2, "traffic": 0.2}. This probability distribution more accurately describes the relative importance of different scene types within the current time window and provides a quantitative basis for subsequent optimization processing.
[0057] By calculating the frequencies of scene types and performing normalization, the present application actually extracts the statistical features of scene types from the time series. This frequency-based probability estimation method is simple and efficient, can make full use of the time series information, and capture the overall distribution characteristics of scene types. At the same time, the normalization process makes the probability distribution comparable and consistent, facilitating the integration with other prior knowledge.
[0058] It should be noted that in practical applications, the present application may encounter the situation where certain scene types do not appear at all in the N-frame sequence. In this case, both the frequency and the normalized probability will be 0, which may cause inconvenience to subsequent calculations. To address this situation, the present application can adopt smoothing techniques, such as Laplace smoothing or Good-Turing estimation, to assign a small initial probability to each scene type to avoid the case of probability being 0.
[0059] S303, according to the preset scene type transition probability matrix, correct the scene type probability to obtain the corrected scene type probability, where the scene type transition probability matrix represents the prior probability of scene type conversion between adjacent frames; Specifically, after obtaining the scene type probability, the present application also needs to use the preset scene type transition probability matrix to correct the scene type probability to obtain the corrected scene type probability. The reason for this correction is that the scene types between adjacent frames are not completely independent, but there is a certain correlation and transition rule. Using this prior knowledge, the present application can further improve the accuracy and coherence of scene type probability estimation.
[0060] The scene type transition probability matrix is a preset matrix that quantifies the prior probability of scene type conversion between adjacent frames. Specifically, each element of the matrix represents the probability of transitioning from one scene type to another. For example, for a transition probability matrix containing three scene types: "indoor", "outdoor", and "traffic", the element P(indoor|outdoor) represents the probability that the next frame transitions to the "indoor" scene when the current frame is the "outdoor" scene.
[0061] This transition probability matrix is usually preset based on a large amount of data statistics and expert experience, reflecting the general transition rules between scene types. For example, the probability of transitioning from "indoor" to "outdoor" is usually relatively high, while the probability of directly transitioning from "indoor" to "traffic" is relatively low.
[0062] With the scene type transition probability matrix, the present application can correct the scene type probability obtained in the previous step. The correction process is actually a process of Bayesian inference, which combines the prior knowledge (transition probability) with the observed data (scene type probability) to obtain the posterior probability (corrected scene type probability).
[0063] Specifically, the correction process can be achieved through matrix multiplication. In this application, the scene type probability vector is multiplied by the transition probability matrix to obtain a new probability vector, which is the corrected scene type probability. Mathematically, this process can be expressed as: P(S_t|O_1:t) = P(O_t|S_t) * P(S_t|S_{t-1}) * P(S_{t-1}|O_1:{t-1}), where S_t represents the scene type at time t, O_t represents the observation at time t (scene type probability), and P(S_t|O_1:t) represents the posterior probability of the scene type given the first t observations.
[0064] Through the correction of the transition probability matrix, this application can make full use of the correlation information between scene types and obtain a more accurate and coherent scene type estimation. The corrected scene type probability not only considers the current observation but also integrates historical information and prior knowledge, making the estimation result smoother and more stable and reducing the occurrence of mutations and discontinuities.
[0065] S304. Based on the corrected scene type probability and the preset Markov prediction model, combined with the historical scene type sequence, predict the scene type probability of the next frame to obtain the predicted scene type probability; Specifically, after obtaining the corrected scene type probability, this application further predicts the scene type probability of the next frame based on this probability and the preset Markov prediction model, combined with the historical scene type sequence, to obtain the predicted scene type probability. The reason for this prediction step is that the change of scene types often has a certain temporal continuity and smoothness. By predicting the future scene, this application can further improve the accuracy and stability of the entire scene type sequence estimation.
[0066] The Markov prediction model is a commonly used time series prediction model that assumes that the future state depends only on the current state and is independent of past states. In the problem of scene type prediction, this application can regard the scene type of each frame as a state in the Markov model and use the scene type probability and transition probability matrix of the current frame to predict the scene type probability of the next frame.
[0067] Specifically, the prediction process can be achieved through the following formula: P(S_{t+1}|O_1:t) = \sum_{S_t} P(S_{t+1}|S_t) * P(S_t|O_1:t), where S_t represents the scene type at time t, O_t represents the observation at time t (scene type probability), P(S_{t+1}|S_t) represents the transition probability of the scene type from time t to t+1, and P(S_t|O_1:t) represents the posterior probability of the scene type at time t, that is, the corrected scene type probability.
[0068] Through the above formula, this application can calculate the prediction probability of each scene type in the next frame (time t+1). This prediction probability comprehensively considers the current scene type estimation (corrected probability) and the transition law between scene types (transition probability matrix), so it can give a more accurate and coherent future scene prediction.
[0069] In actual calculation, this application can also use the historical scene type sequence to further optimize the prediction process. Specifically, this application can use the scene type estimation results in the past period (such as a few seconds) as historical information and incorporate it into the current prediction in the way of a time window or a decay factor. In this way, the prediction result not only considers the current state but also takes into account a certain historical context, making the prediction more stable and smooth.
[0070] The predicted scene type probability obtained through the Markov prediction model provides a reasonable expectation for the future scene change of this application. This expectation can not only be used for the optimization of the scene type but also provide guidance for the subsequent AR content generation and interaction strategy. For example, when the prediction result shows that an outdoor scene may appear in the future, the AR system can adjust the display parameters in advance to adapt to the lighting conditions of the outdoor environment; when the prediction result shows that it may enter a traffic scene, the AR system can load relevant navigation information and road condition tips in advance to provide more timely and accurate services for users.
[0071] S305, perform weighted fusion on the corrected scene type probability and the predicted scene type probability to obtain a fusion probability; Specifically, after obtaining the corrected scene type probability and the predicted scene type probability, this application needs to perform weighted fusion on these two probabilities to obtain the final fusion probability. The reason for this step of fusion is that the corrected scene type probability and the predicted scene type probability respectively reflect two important aspects of the scene type estimation: current observation and future expectation. Through weighted fusion, this application can comprehensively consider the information of these two aspects to obtain a more accurate and reliable scene type estimation.
[0072] Specifically, weighted fusion can be achieved through the following formula: P(S_t) = w_1 * P(S_t|O_1:t) + w_2 * P(S_{t + 1}|O_1:t), where P(S_t) represents the fusion probability of the scene type at time t, P(S_t|O_1:t) represents the posterior probability (corrected probability) of the scene type at time t, P(S_{t + 1}|O_1:t) represents the predicted probability of the scene type at time t + 1, and w_1 and w_2 are the weight coefficients of the two probabilities, satisfying w_1 + w_2 = 1.
[0073] The values of the weight coefficients w_1 and w_2 reflect the relative importance attached by this application to the current observation and future expectation. Generally speaking, w_1 should be greater than w_2 because the current observation is more direct and reliable, while the future expectation has certain uncertainties. However, the specific weight values need to be adjusted according to the actual application scenario and data characteristics. For example, in a scenario with rapid changes, this application may wish to increase the weight of w_2 to consider more of the future expectation; while in a scenario with slow changes, this application can increase the weight of w_1 and rely more on the current observation.
[0074] Through weighted fusion, this application obtains a scene type probability that comprehensively considers the current observation and future expectation. This fusion probability is more stable and reliable than using only the corrected probability or the predicted probability alone, and can effectively balance the information in the time dimension and improve the overall accuracy of scene type estimation.
[0075] S306. According to the preset scene type determination rule, perform threshold screening on the fusion probability to obtain the target scene type.
[0076] Specifically, after obtaining the fusion probability, this application also needs to perform threshold screening on the fusion probability according to the preset scene type determination rule to obtain the final target scene type. The reason for this screening step is that although the fusion probability has comprehensively considered various aspects of information, there may still be some uncertainties and noises in it. Through threshold screening, this application can further improve the reliability of scene type estimation and obtain a target scene type with high confidence.
[0077] The scene type determination rule is a set of pre - set rules used to map the fusion probability to the final scene type. These rules are usually derived from prior knowledge and experience summary, reflecting the importance and distinctiveness of different scene types in practical applications. For example, this application can set a threshold. Only when the fusion probability of a certain scene type exceeds this threshold, it is regarded as a possible target scene type; and when the fusion probabilities of all scene types are lower than the threshold, this application may need to introduce an "unknown" or "other" type to indicate that the current scene cannot be determined.
[0078] Specifically, the process of threshold screening can be achieved through the following steps: First, this application sorts the fusion probabilities and finds the scene type with the largest probability value. Then, this application determines whether the probability value of this scene type exceeds the preset threshold. If it exceeds the threshold, this application determines this scene type as the target scene type; if it does not exceed the threshold, this application can consider the scene type with the second - largest probability and determine whether it exceeds the threshold, and so on, until a scene type that exceeds the threshold is found, or all scene types do not meet the conditions.
[0079] After finding the scene type that meets the threshold condition, this application can further consider other determination rules. For example, this application can set a minimum duration. Only when a certain scene type continuously meets the threshold condition for a period of time, it is determined as the target scene type. This can avoid misjudgment caused by short - time noise or mutations. In addition, this application can also consider the semantic relationship between scene types, such as the mutual exclusivity between indoor scenes and outdoor scenes, to further improve the accuracy of determination.
[0080] Through threshold screening and scene type determination rules, this application finally obtains a target scene type with high confidence. This target scene type comprehensively considers various factors such as current observations, historical information, and future expectations. After multiple screenings and optimizations, it has high reliability and stability. It can accurately reflect the environmental semantics where the user is currently located and provide a reliable basis for subsequent AR content generation and interaction.
[0081] Based on the above - mentioned embodiments, as an optional embodiment, matching the corresponding application scene template from the preset scene template library according to the scene type includes: S401, constructing a knowledge graph in the form of triples for the scene templates in the preset scene template library. The triples include scene type, scene attribute, and scene parameter, to obtain a semantic network structure; Specifically, after obtaining the target scene type, the present application needs to match the corresponding application scene template from a preset scene template library. To achieve efficient and accurate template matching, the present application first constructs a knowledge graph in the form of triples for the scene templates in the preset scene template library to obtain a semantic network structure. The reason for constructing the knowledge graph is that the knowledge graph can represent the semantic relationships between scene templates in a structured manner, facilitating subsequent queries and matching.
[0082] Specifically, the present application represents each scene template as a triple, including three elements: scene type, scene attribute, and scene parameter. The scene type represents the scene category to which the template belongs, such as indoor, outdoor, traffic, etc.; the scene attribute represents the feature description of the template, such as lighting conditions, spatial layout, environmental semantics, etc.; the scene parameter represents the numerical parameters of the template, such as color, texture, geometric dimensions, etc. Through such triple representation, the present application can transform the scene template library into a structured knowledge base.
[0083] When constructing the knowledge graph, the present application not only needs to represent the internal structure of each scene template, but also depict the semantic relationships between different templates. These semantic relationships can be based on the hierarchical structure of scene types (such as offices, living rooms, etc. in indoor scenes), or on the similarity of scene attributes (such as scenes with similar lighting conditions), or on the numerical relationships of scene parameters (such as the proximity of colors and sizes), etc. Through these semantic relationships, the present application can organize the scattered scene templates into an interconnected network structure to form a complete semantic graph.
[0084] After constructing the knowledge graph, the present application obtains a representation of the scene template library with rich semantics and clear structure. Compared with the original flat template library, the template library in the form of a knowledge graph has the following advantages: First, it organizes and associates the relationships between templates in a semantic form, facilitating semantic-based queries and inferences; second, it provides multi-dimensional and multi-granularity template representations, which can be matched from different perspectives such as scene type, attribute, and parameter; finally, it transforms the template matching problem into a search problem on a graph, and can use graph algorithms and machine learning techniques for efficient solution.
[0085] S402, extract the semantic features of the scene type, generate the feature vector of the scene type, and generate the corresponding template vector based on the attribute information of each template node in the knowledge graph; Specifically, after constructing the scenario template knowledge graph, in order to achieve the accurate matching between the target scenario type and the template nodes, this application needs to extract the semantic features of the scenario type, generate the feature vector of the scenario type, and generate the corresponding template vector based on the attribute information of each template node in the knowledge graph. The reason for generating these two types of vectors is that the vector representation can map the scenario type and the template nodes into the same semantic space, facilitating subsequent matching through vector similarity calculation.
[0086] Specifically, for the generation of the feature vector of the scenario type, this application can use natural language processing techniques, such as word embedding, semantic parsing, etc., to map the scenario type text into a real-valued vector with a fixed dimension. Each element in this vector represents the weight of the scenario type in different semantic dimensions, reflecting the semantic attributes of the scenario type. For example, for the scenario type of "outdoor park", its feature vector may have relatively high weight values in dimensions such as "natural environment", "open space", and "leisure activities".
[0087] Similarly, for the generation of the vector of the template node, this application can extract its semantic features based on the attribute information of this node in the knowledge graph. These attribute information include the type, relationship, and associated entities of the node. For example, for a template node of "indoor office", this application can extract features from its associated attributes such as "lighting conditions", "furniture layout", and "office equipment" to generate a vector that comprehensively describes the semantics of this template node.
[0088] When generating the feature vector, this application can also use the structural information of the knowledge graph, such as the neighbors and paths of the nodes, to enhance the expression ability of the features. For example, this application can consider indicators such as the centrality and connectivity of the nodes in the graph to reflect their importance in the entire scenario system; this application can also use the semantic relationships between the nodes, such as synonymy and hyponymy, to expand and enrich the feature representation of the nodes.
[0089] Through feature vectorization, this application transforms the originally heterogeneous and unstructured scenario types and template nodes into a unified vector representation form. This vector representation is not only concise and efficient but also has good semantic expression ability, which can depict the characteristics of the scenario type and the template nodes in different semantic dimensions. At the same time, vectorization also provides convenience for subsequent similarity calculation, clustering, sorting, etc., enabling this application to efficiently find the template node that best matches the target scenario type in the massive template library.
[0090] S403, calculate the cosine similarity between the feature vector of the scenario type and the template vector to obtain the similarity distribution; Specifically, after obtaining the feature vector and template vector of the scene type, in order to find the template node that best matches the target scene type, this application needs to calculate the similarity between the feature vector of the scene type and the template vector to obtain a similarity distribution. The reason for using similarity calculation is that similarity can quantitatively describe the proximity of two vectors in the semantic space. The more similar the vectors are, the closer the corresponding scene types and template nodes are semantically.
[0091] Specifically, this application selects cosine similarity as the metric standard for similarity. Cosine similarity is a commonly used method for calculating vector similarity. It measures the similarity between two vectors by calculating the cosine value of the angle between them. The value range of cosine similarity is between [-1, 1]. The larger the value, the more similar the two vectors are, and the smaller the value, the less similar the two vectors are.
[0092] The formula for calculating cosine similarity is: cos(θ) = (A·B) / (||A||·||B||), where A and B are two vectors respectively, A·B represents the dot product of the vectors, ||A|| and ||B|| represent the magnitudes of the vectors, and θ represents the angle between the two vectors.
[0093] In actual calculation, this application calculates the cosine similarity between the feature vector of the scene type and the vector of each template node in the knowledge graph in turn to obtain a similarity value. This process can be implemented through batch matrix operations to improve calculation efficiency. After the calculation is completed, this application obtains a similarity distribution, where each value represents the similarity degree between the target scene type and the corresponding template node.
[0094] Through the similarity distribution, this application can intuitively understand the matching degree between the target scene type and different template nodes. The template nodes with larger values in the similarity distribution indicate that they are closer to the target scene type in semantic features and are more likely to be the matching results; while the template nodes with smaller similarity values indicate that they are quite different from the target scene type and are less likely to be the best match.
[0095] The similarity distribution provides a quantitative and continuous representation of the matching degree, enabling this application to rank, filter, and optimize the matching results in the entire template space. At the same time, the similarity distribution also provides an important basis for recommendation systems, semantic search, etc., enabling this application to recommend the most relevant and matching scene templates based on the user's scene type preferences.
[0096] S404, normalize the similarity distribution through the softmax function to obtain the matching probability distribution of each template node; Specifically, after obtaining the similarity distribution between the scene type and the template nodes, in order to further quantify the confidence level of the matching, this application needs to normalize the similarity distribution to obtain the matching probability distribution of each template node. The reason for normalization is that although the original similarity values can reflect the relative priorities of the matching, they are not a strict probability measure and cannot intuitively explain the confidence of the matching. Through normalization, this application can map the similarity values to the interval [0, 1], making them satisfy the basic properties of probability and being more convenient for subsequent decision-making and applications.
[0097] Specifically, this application uses the softmax function to normalize the similarity distribution. The Softmax function is a commonly used normalization method, especially suitable for calculating and comparing probabilities in multi-classification problems. Its basic idea is to map a set of real numbers to a set of probability values whose sum is 1, while preserving the relative magnitude relationship of the original numerical values.
[0098] The mathematical definition of the Softmax function is: P(i) = exp(xi) / ∑exp(xj), where xi represents the i-th real number, P(i) represents the i-th probability value after normalization, and ∑exp(xj) represents the normalization factor of the sum of the exponents of all real numbers.
[0099] In the scene matching problem of this application, this application regards the similarity values between the scene type and each template node as the input of the softmax function, and calculates the probability value after normalization through the above formula. This process can be efficiently implemented through vectorized operations. After normalization, this application obtains a matching probability distribution, where each probability value represents the matching confidence level between the target scene type and the corresponding template node.
[0100] Through softmax normalization, the originally heterogeneous and unbounded similarity values are mapped to a unified probability scale. The matching probability distribution has the following advantages: First, its value range is between [0, 1], which conforms to the basic constraints of probability and is convenient for interpretation and application; Second, it preserves the relative magnitude relationship of the original similarity values, and the larger the probability value, the higher the matching confidence level; Finally, it normalizes the matching results of different template nodes to the same scale, which is convenient for comparison and sorting between different scene types.
[0101] S405, screen the matching probability distribution based on a preset matching threshold, select the template node with the highest matching probability and greater than the matching threshold as the target template node, and obtain the corresponding scene attribute information in the knowledge graph based on the target template node; Specifically, after obtaining the normalized matching probability distribution, in order to determine the final target template from the candidate template nodes, this application needs to screen the matching probability distribution based on a preset matching threshold. The reason for setting the matching threshold is that although the matching probability distribution can reflect the relative matching degree of different template nodes, not all candidate nodes are suitable matching targets. To ensure the quality and usability of the matching result, this application needs to set a threshold to filter out the nodes with too low matching confidence and only retain the target templates with sufficiently high confidence.
[0102] Specifically, the setting of the matching threshold needs to comprehensively consider various factors, such as application requirements, data characteristics, user experience, etc. Generally speaking, the higher the threshold is set, the higher the matching accuracy will be, but the recall rate may decrease; while the lower the threshold is set, the higher the recall rate of the matching will be, but the accuracy may decline. This application needs to find a suitable balance between accuracy and recall rate to meet the needs of actual applications.
[0103] After determining the matching threshold, this application compares each probability value in the matching probability distribution with the threshold. For those template nodes whose probability values are greater than or equal to the threshold, this application regards them as potential matching targets; while for those template nodes whose probability values are less than the threshold, this application excludes them from the candidate set. Through this screening process, this application can obtain a set of target template nodes with relatively high confidence.
[0104] In actual operation, this application can further optimize the screening result. For example, this application can select the template node with the highest matching probability from the candidate set as the final target template. This can ensure that on the basis of the highest semantic similarity, the result with the highest confidence is preferentially matched. In addition, if there are multiple template nodes with similar confidence levels in the candidate set, this application can also introduce other factors, such as user preferences, context information, etc., to assist in sorting and selection.
[0105] After determining the target template node, this application also needs to obtain the corresponding scene attribute information in the knowledge graph based on this node. The scene attribute information is a more specific and fine-grained description of the target scene template, including multiple aspects such as the environmental parameters, interaction logic, and rendering strategy of the scene. By obtaining this attribute information, this application can provide necessary reference and guidance for the subsequent generation and presentation of AR content.
[0106] In the knowledge graph, each template node is connected to a set of attribute nodes, and these attribute nodes depict various features and parameters of this template. Therefore, the process of obtaining scene attribute information is essentially a process of performing associative queries and information aggregation in the knowledge graph. This application can start from the target template node and extract the complete set of attribute information through the attribute edges connected to its associated attribute nodes.
[0107] The acquisition of scene attribute information enables the AR system to have a more comprehensive and detailed understanding of the target scene. With these attribute parameters, the AR system can dynamically adjust the content presentation method according to the actual environmental conditions to provide a more natural and immersive user experience. For example, according to the lighting attribute, the system can adjust the brightness and contrast of the AR content in real time; according to the spatial layout attribute, the system can optimize the size and position of the AR content; according to the interaction logic attribute, the system can provide a more natural and smooth interaction method.
[0108] S406. Determine the corresponding application scenario template according to the scene attribute information.
[0109] Specifically, after obtaining the scene attribute information of the target template node, the present application needs to finally determine the application scenario template corresponding to the current scene type according to these attribute information. The reason for this step is that the scene attribute information provides a more specific and fine-grained description of the target scene, including various parameters and settings required for constructing and presenting AR content. By comprehensively using these attribute information, the present application can select the best from the candidate template nodes to generate the final application scenario template to guide the subsequent generation and interaction of AR content.
[0110] Specifically, the process of determining the application scenario template can be regarded as a template reconstruction process based on attribute matching. The present application first extracts the names and values of each attribute from the attribute information obtained from the target template node. These attributes can cover all aspects of the scene, such as environmental parameters (such as lighting, weather), spatial layout (such as size, position), interaction logic (such as event triggering, feedback mechanism), rendering strategy (such as material, texture), etc.
[0111] Then, the present application matches the extracted attribute names and values with the predefined application scenario template. The predefined application scenario template is a parameterized scene description framework, which contains a series of placeholders and default values. These placeholders and default values correspond to different scene attributes and constitute a complete scene description. In the matching process, the present application replaces the corresponding placeholders in the application scenario template with the attribute values of the target template node and sets the unspecified attributes to default values. In this way, the present application can obtain an application scenario template customized for the current scene.
[0112] When performing attribute matching, the present application can also introduce some intelligent optimization mechanisms to improve the efficiency and quality of template generation. For example, the present application can pre-statistically analyze the high-frequency attributes of different scenario types to construct an attribute priority table; during the matching process, the present application preferentially matches the attributes with higher priorities to quickly determine the main features of the scenario. Another example is that the present application can establish an association rule library between attributes to describe the constraints and dependencies between different attributes; when generating a template, the present application checks whether the attribute values satisfy these association rules to ensure the rationality and consistency of the scenario description.
[0113] Through the above steps, the present application finally determines the application scenario template corresponding to the current scenario type. This template integrates the scenario semantic information and specific attribute parameters in the knowledge graph, providing direct input and basis for subsequent AR content generation. Based on the application scenario template, the AR system can quickly and automatically construct virtual content that highly matches the real scenario and provide natural and smooth user interaction according to the interaction logic.
[0114] The introduction of the application scenario template enables the AR system to have the ability to quickly adapt to different scenarios. For each identified target scenario type, the system can use the corresponding scenario attributes to dynamically generate the most suitable AR content template. This template-based content generation mechanism not only improves the development efficiency of AR applications and reduces the cost of content production; more importantly, it endows the AR system with the ability of environmental understanding and autonomous learning, enabling AR technology to truly integrate into people's daily life and work and providing a seamless and natural human-computer interaction experience.
[0115] S103. Determine the types of environmental parameters to be collected according to the application scenario template, and collect environmental data through the sensors corresponding to the types of environmental parameters to be collected to obtain target environmental data; Specifically, first, based on the target scenario type obtained in the previous step, match the corresponding application scenario template from the scenario template library. The application scenario template is essentially a semantic description of environmental parameter requirements, which predefines the key environmental factors that need to be concerned in a specific scenario. For example, for the outdoor water source detection scenario, the application scenario template will guide the selection of specific environmental parameter types such as temperature, humidity, light intensity, and water quality indicators.
[0116] After determining the application scenario template corresponding to the target scenario type, the system will accurately match the corresponding sensors according to the preset environmental parameter types in the template. This process is not a simple mechanical correspondence, but an accurate mapping achieved through semantic associations in the knowledge graph. Specifically, the system will extract the semantic features of the scenario type, generate feature vectors, and calculate the cosine similarity with the template vectors generated from the attribute information of each template node in the scenario template library, and finally select the template node with the highest matching degree.
[0117] S104, generate AR device adjustment parameters according to the target environmental data; Specifically, in the process of generating AR device adjustment parameters from the target environmental data, the core problem that the system first faces is how to convert complex environmental information into operable device parameters. The limitations of traditional methods that rely only on a single light sensor cannot comprehensively perceive the multi-dimensional characteristics of the environment, resulting in the AR display effect being difficult to adapt to complex environments. Therefore, the system designs a parameter generation method based on the digital twin model.
[0118] Through in-depth analysis of the target environmental data, the system constructs a digital twin model that includes a three-dimensional geometric model and a physical property model. This model is not just a static description of the environment, but an intelligent model that can dynamically simulate and predict environmental characteristics. In the construction process, the system performs semantic reconstruction and mathematical modeling on multi-dimensional information such as the spatial structure, optical characteristics, temperature and humidity of the environment.
[0119] The three-dimensional geometric model accurately restores the spatial structure characteristics of the environment, including object contours, spatial positions, and surface textures. To obtain more accurate environmental occlusion information, the system uses a ray tracing algorithm to calculate the environmental light occlusion factor. Ray tracing is a high-precision calculation method for simulating the propagation of light in three-dimensional space. By tracing the paths of multiple light rays emitted from the light source, the system can accurately calculate the proportion of blocked light.
[0120] In the ray tracing calculation, the system not only traces direct light rays, but also simulates the processes of light reflection, refraction, and scattering. This fine simulation of light propagation makes the calculation of the environmental light occlusion factor more accurate and can comprehensively reflect the complex influence of the environment on light propagation. By analyzing the occlusion effect of each geometric body on light, the system establishes a dynamic optical environment model.
[0121] Based on the ambient occlusion factor and ambient light intensity, the system calculates the transparency parameter τ of the AR overlay content. The calculation of the transparency parameter follows the formula: τ = k·(Lmax + Lambient) / Lambient, where k is the global adjustment coefficient, Lambient is the ambient light intensity, and Lmax is the maximum brightness of the AR display device. This formula adjusts the transparency of the AR content to ensure the best visibility of virtual information under different lighting conditions.
[0122] Based on the above embodiments, as an alternative embodiment, generating AR device adjustment parameters according to target environment data includes: S501, constructing a digital twin model according to the target environment data, where the digital twin model includes a joint representation of a three-dimensional geometric model and a physical property model; S502, calculating the ambient occlusion factor through a ray tracing algorithm based on the ambient occlusion relationship in the three-dimensional geometric model; S503, calculating the transparency parameter τ of the AR overlay content according to the ambient occlusion factor and ambient light intensity, where: τ = k·(Lmax + Lambient) / Lambient, where k is the global adjustment coefficient, Lambient is the ambient light intensity, and Lmax is the maximum brightness of the AR display device; S504, calculating the contrast adjustment coefficient of the AR display device based on environmental parameters such as temperature and humidity in the physical property model; S505, generating AR device adjustment parameters according to the transparency parameter and the contrast adjustment coefficient.
[0123] Specifically, in step S501, the system constructs a digital twin model according to the target environment data. The digital twin model is a modeling method that precisely maps a physical entity in the digital space. It constructs a joint representation including a three-dimensional geometric model and a physical property model through the multi-dimensional information of the target environment data. During the construction process, the system first performs semantic parsing and structural reconstruction on the collected environment data, converting the scattered environment information into a computable digital model.
[0124] The three-dimensional geometric model accurately restores the spatial structure characteristics of the environment, including geometric information such as object contours, spatial positions, and surface textures. The physical property model embeds environmental parameters such as temperature, humidity, and light intensity, establishing a digital description of the environmental physical characteristics. This joint representation enables the digital twin model not only to visualize the environment but also to perform complex physical property simulations and predictions.
[0125] In step S502, the system analyzes the environmental occlusion relationship based on the three-dimensional geometric model and calculates the ambient light occlusion factor through the ray tracing algorithm. Ray tracing is an accurate calculation method that simulates the propagation of light in three-dimensional space and can simulate the complex interactions of light with various objects in the scene. The system calculates the proportion of occluded light by tracing multiple light ray paths emitted from the light source, quantifying the degree of interference of the environment on light propagation.
[0126] The calculation of the ambient light occlusion factor takes into account the influence of each geometric body in the scene on light propagation. This is not just a simple geometric occlusion calculation but an accurate physical simulation of the behavior of light propagation in a complex environment. By analyzing the reflection, refraction, and occlusion of light, the system can accurately evaluate the impact of the environment on light propagation.
[0127] In step S503, the system calculates the transparency parameter τ of the AR overlay content based on the ambient light occlusion factor and the ambient light intensity. The calculation formula for the transparency parameter is τ = k·(Lmax + Lambient) / Lambient, where k is the global adjustment coefficient, Lambient is the ambient light intensity, and Lmax is the maximum brightness of the AR display device. This formula dynamically adjusts the transparency of the AR content to ensure the clear visibility of virtual information under different lighting conditions.
[0128] When the ambient light intensity is low, the value of τ increases, and the AR content will be more opaque, ensuring the readability of information; in a strong light environment, the value of τ decreases, and the AR content will be more transparent, avoiding excessive occlusion of the real scene. The global adjustment coefficient k provides an additional fine-tuning mechanism to further optimize the adaptability of the transparency parameter.
[0129] In step S504, the system calculates the contrast adjustment coefficient of the AR display device based on environmental parameters such as temperature and humidity in the physical property model. Temperature may cause optical components to expand and deform, and humidity may cause condensation on the optical surface. These factors will significantly affect the optical performance of the display device. The contrast adjustment coefficient accurately predicts and compensates for the impact of these environmental factors on the display quality by establishing a mathematical model between environmental parameters and display performance degradation.
[0130] Finally, in step S505, the system generates the AR device adjustment parameters based on the transparency parameter and the contrast adjustment coefficient. These parameters comprehensively cover the key performance indicators of the display system, including brightness, contrast, color saturation, and image sharpness. The generation process uses a multi-objective optimization algorithm to find the best balance between ensuring display clarity, user visual comfort, and device performance.
[0131] S105. Adjust the display parameters of the AR device according to the AR device adjustment parameters.
[0132] Specifically, the adjustment of display parameters of traditional AR devices often relies on simple linear mapping of a single sensor and cannot meet the fine display requirements in complex environments. In this embodiment, through systematic parameter mapping and precise regulation, the key problem of insufficient AR display adaptability is solved.
[0133] The core of the adjustment of display parameters of the AR device lies in accurately converting the adjustment parameters generated in the previous steps into display control instructions that can be actually executed. The system first establishes a mapping relationship between the adjustment parameters and the display hardware parameters. This mapping is not a simple linear correspondence, but is based on a complex non-linear conversion model.
[0134] In specific implementation, the system first deconstructs the parameters of the display hardware of the AR device. The display parameters include but are not limited to multiple dimensions such as brightness, contrast, color saturation, color temperature, image sharpness, etc. Each parameter corresponds to a specific hardware control interface, such as voltage control of the liquid crystal panel, backlight brightness adjustment, parameter configuration of the image processing chip, etc.
[0135] The system accurately controls the transparency of the AR content according to the transparency parameter τ generated in the previous steps. When the value of τ is large, it means that the ambient light is weak, and the system will increase the opacity of the AR content to improve the visibility of virtual information; when the value of τ is small, the system reduces the opacity of the AR content to ensure the clarity of the real scene. This dynamic adjustment avoids the problem that virtual content is too obvious or too dim in traditional methods.
[0136] The introduction of the contrast adjustment coefficient further optimizes the display effect. The system accurately compensates for the performance attenuation of the optical system based on the contrast adjustment coefficient calculated from parameters such as ambient temperature and humidity. For example, in a high-temperature and high-humidity environment, the optical components may undergo slight deformation, resulting in a decrease in contrast. Through the contrast adjustment coefficient, the system can compensate for this performance degradation in real time and maintain the stability of the display quality.
[0137] Another key link in the adjustment of display parameters is color and image processing. The system dynamically adjusts the color balance, saturation, and sharpness of the image through the image processing chip according to the ambient light and occlusion factor. This adjustment is not just a simple linear scaling of parameters, but is based on perceptual psychology and image processing algorithms to ensure the best visual experience in different environments.
[0138] Based on the above embodiments, as an alternative embodiment, warning information is generated according to the target environment data and preset risk assessment rules, including: S106, obtain the environmental parameter values in the target environment data, compare the environmental parameter values with the preset risk thresholds, and determine the risk level when the environmental parameter values exceed the risk thresholds; S107, generate corresponding warning information according to the risk level.
[0139] Specifically, traditional environmental monitoring methods usually rely on single - parameter threshold judgment and cannot comprehensively evaluate potential risks in complex environments. In this embodiment, a multi - dimensional and intelligent environmental risk assessment and early warning mechanism is designed to improve the accuracy and timeliness of environmental perception.
[0140] In step S106, the system first extracts key environmental parameter values from the target environmental data. The risk assessment rule is a pre - defined environmental safety judgment mechanism that converts complex environmental parameters into quantifiable risk indicators. The system precisely compares the extracted environmental parameter values, such as temperature, humidity, light intensity, air quality, etc., with the preset risk thresholds.
[0141] The construction of risk thresholds is based on long - term environmental monitoring data and professional domain knowledge, and is systematically established through statistical analysis and machine learning methods. For each environmental parameter, the system pre - sets multiple risk thresholds, and these thresholds divide different risk levels. For example, the risk threshold for the temperature parameter may be set as: 20 - 30 degrees Celsius is the safe range, and values below 20 or above 30 enter different levels of risk.
[0142] The comparison process uses a multi - dimensional and multi - level risk judgment algorithm. The system not only checks whether a single parameter exceeds the standard, but also comprehensively analyzes the correlation and mutual influence between parameters. If a certain environmental parameter value exceeds the preset risk threshold, the system will accurately evaluate the risk level according to the severity of the exceedance and the importance of the parameter. The risk level is usually divided into three levels: low risk, medium risk, and high risk.
[0143] The determination of the risk level is not a simple threshold judgment, but is based on a complex multi - factor weight model. The system considers multiple dimensions such as the amplitude of parameter exceedance, duration, and the importance of parameters. For example, a slight temperature exceedance may be judged as low risk, but if accompanied by abnormalities in other parameters at the same time, it may be upgraded to medium or high risk.
[0144] In step S107, the system generates corresponding early warning information based on the risk level determined in the previous step. The generation of early warning information follows three key principles: accuracy, timeliness, and operability. For the low - risk level, the early warning information is presented in a prompt manner, emphasizing a slight environmental anomaly; for the medium - risk level, the early warning information more clearly points out potential dangers and gives preliminary suggestions; for the high - risk level, the early warning information is prominently displayed in a warning form, accompanied by detailed risk descriptions and coping strategies.
[0145] The semantic generation of warning information uses natural language processing technology to ensure that the information is both accurate and easy to understand. The built-in risk description template library in the system supports generating accurate and clear descriptions for different types and levels of risks. Warning information is not just a simple presentation of data, but an intelligent interpretation and prediction of environmental risks.
[0146] Based on the above embodiments, as an optional embodiment, warning information is generated according to the target environmental data and preset risk assessment rules, including: S601, Obtain the environmental parameter values in the target environmental data, and compare the environmental parameter values with the preset risk thresholds. When the environmental parameter values exceed the risk thresholds, determine the risk level; S602, Generate corresponding warning information according to the risk level.
[0147] Specifically, traditional environmental monitoring methods have prominent problems such as single risk identification and inaccurate warning. This embodiment designs a multi-dimensional and intelligent environmental risk assessment mechanism to improve the accuracy and timeliness of environmental risk perception.
[0148] In step S601, the system extracts the key environmental parameter values from the target environmental data. The risk assessment rule is a predefined environmental safety judgment mechanism that converts complex environmental parameters into quantifiable risk indicators. The system precisely compares the extracted environmental parameter values, such as temperature, humidity, light intensity, etc., with the preset risk thresholds.
[0149] The construction of the risk thresholds is based on long-term environmental monitoring data and professional domain knowledge, and is systematically established through statistical analysis and machine learning methods. For each environmental parameter, the system has preset multiple risk thresholds, and these thresholds divide different risk levels. For example, the risk threshold for the temperature parameter may be set as: 20 - 30 degrees Celsius is the safe range, and below 20 or above 30 enters different levels of risk.
[0150] The comparison process uses a multi-dimensional and multi-level risk judgment algorithm. The system not only checks whether a single parameter exceeds the standard, but also comprehensively analyzes the relevance and mutual influence between parameters. The determination of the risk level considers multiple dimensions such as the extent of parameter over-standard, duration, and parameter importance. Specifically, the system establishes a multi-factor weight model and dynamically calculates the risk level according to the deviation degree and importance of environmental parameters.
[0151] When a certain environmental parameter value exceeds the preset risk threshold, the system will accurately evaluate the risk level according to the severity of the over-standard and the importance of the parameter. The risk level is usually divided into three levels: low risk, medium risk, and high risk. For example, a slight over-standard of temperature may be judged as low risk, but if accompanied by abnormalities in other parameters at the same time, it may be upgraded to medium or high risk.
[0152] In step S602, the system generates corresponding warning information based on the risk level determined in the previous step. The generation of the warning information follows three key principles: accuracy, timeliness, and operability. For a low risk level, the warning information is presented in a prompt manner, emphasizing a slight abnormality in the environment; for a medium risk level, the warning information more clearly points out potential dangers and gives preliminary suggestions; for a high risk level, the warning information is highlighted in the form of a warning, accompanied by a detailed risk description and coping strategies.
[0153] The semantic generation of the warning information uses natural language processing technology to ensure that the information is both accurate and easy to understand. The built-in risk description template library in the system supports generating accurate and clear descriptions for different types and levels of risks. The warning information is not just a simple presentation of data, but an intelligent interpretation and prediction of environmental risks.
[0154] On the other hand, the present application also provides an environment-based perception enhancement system, such as Figure 2 , and the system includes: A data acquisition module 1, configured to obtain N environmental images through the camera of the AR device, obtain an environmental image set, and perform scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type; A template matching module 2, configured to perform optimization processing on the preliminary scene type to obtain a target scene type, and match a corresponding application scene template from a preset scene template library according to the scene type; An environmental parameter determination module 3, configured to determine the type of environmental parameters to be collected according to the application scene template, and collect environmental data through sensors corresponding to the type of environmental parameters to be collected to obtain target environmental data; An adjustment parameter determination module 4, configured to generate AR device adjustment parameters according to the target environmental data; An adjustment module 5, configured to adjust the display parameters of the AR device according to the AR device adjustment parameters.
[0155] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0156] Among them, the communication bus 302 is used to realize the connection and communication between these components.
[0157] Among them, the user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.
[0158] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0159] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling data stored in the memory 305, the processor 301 executes various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 301 may integrate one or a combination of several of the following: a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, the user interface, and application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately through a single chip.
[0160] Among them, the memory 305 may include a random access memory (RAM) and may also include a read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), and instructions for implementing the above-mentioned method embodiments; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 305 may also be at least one storage system located far from the aforementioned processor 301. Refer to Figure 3, in a memory 305 as a computer storage medium, an operating system, a network communication module, a user interface module, and an application program of a flocculant addition analysis method may be included.
[0161] In Figure 3 In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call the application program of the road evaluation method stored in the memory 305. When executed by one or more processors 301, the electronic device 300 is caused to execute the method as described in one or more of the above embodiments. It should be noted that, for the foregoing method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] In several implementation manners provided by the present application, it should be understood that the disclosed system can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed coupling or direct coupling or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the system or unit can be in an electrical or other form.
[0163] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0164] The embodiments of the present application also provide a computer storage medium. The computer storage medium can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the environment-based perception enhancement method as described in the above Figure 1 shown embodiments. The specific execution process can refer to the specific description of the Figure 1 shown embodiments, and will not be elaborated here.
[0165] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0166] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a memory and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0167] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the practice of the present disclosure, those skilled in the art will easily think of other implementation manners of the present disclosure.
[0168] The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, which follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. An environment-based perception enhancement method, characterized in that, The method is applied to an AR device, and the method includes: Obtaining N environmental images through a camera of the AR device to obtain an environmental image set, and performing scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type; Performing optimization processing on the preliminary scene type to obtain a target scene type, and matching a corresponding application scene template from a preset scene template library according to the scene type; Determining the type of environmental parameters to be collected according to the application scene template, and collecting environmental data through a sensor corresponding to the type of environmental parameters to be collected to obtain target environmental data; Generating AR device adjustment parameters according to the target environmental data; Adjusting the display parameters of the AR device according to the AR device adjustment parameters.
2. The method according to claim 1, characterized in that The performing scene classification based on the environmental images in the environmental image set to obtain a preliminary scene type includes: Performing size normalization processing on the environmental images in the environmental image set to obtain standardized images of a preset size, and performing illumination compensation and noise filtering on the standardized images to obtain enhanced images; Using a pre-trained deep convolutional neural network to extract multi-scale feature maps of the enhanced images, and superimposing spatial attention weights on the multi-scale feature maps to obtain enhanced feature maps; Performing feature dimensionality reduction on the enhanced feature maps to obtain feature vectors, and inputting the feature vectors into a preset classifier to obtain prediction probabilities of each scene type; Performing softmax normalization processing on the prediction probabilities to obtain a probability distribution, and screening the probability distribution according to a preset confidence threshold to obtain the preliminary scene type.
3. The method according to claim 1, wherein The performing optimization processing on the preliminary scene type to obtain a target scene type includes: Determining the preliminary scene types of N consecutive environmental images in the environmental image set based on the preliminary scene type, where N is a preset positive integer greater than 1; Calculating the frequency of occurrence of each scene type in the N environmental images, and performing normalization processing on the frequency to obtain scene type probabilities; According to a preset scene type transition probability matrix, correcting the scene type probabilities to obtain corrected scene type probabilities, where the scene type transition probability matrix represents the prior probability of scene type conversion between adjacent frames; Based on the corrected scene type probabilities and a preset Markov prediction model, combining with a historical scene type sequence, predicting the scene type probabilities of the next frame to obtain predicted scene type probabilities; Performing weighted fusion on the corrected scene type probabilities and the predicted scene type probabilities to obtain a fusion probability; According to a preset scene type determination rule, performing threshold screening on the fusion probability to obtain the target scene type.
4. The method according to claim 1, wherein The matching a corresponding application scene template from a preset scene template library according to the scene type includes: Constructing a knowledge graph in the form of triples for the scene templates in the preset scene template library, where the triples include scene types, scene attributes, and scene parameters, to obtain a semantic network structure; Extract the semantic features of the scene type, generate the feature vector of the scene type, and generate the corresponding template vector based on the attribute information of each template node in the knowledge graph; Calculate the cosine similarity between the feature vector of the scene type and the template vector to obtain the similarity distribution; Perform normalization processing on the similarity distribution through the softmax function to obtain the matching probability distribution of each template node; Based on a preset matching threshold, screen the matching probability distribution, select the template node with the highest matching probability and greater than the matching threshold as the target template node, and obtain the corresponding scene attribute information in the knowledge graph based on the target template node; Determine the corresponding application scene template according to the scene attribute information.
5. The method according to claim 1, wherein The generating the AR device adjustment parameters according to the target environmental data includes: Construct a digital twin model according to the target environmental data, where the digital twin model includes a joint representation of a three-dimensional geometric model and a physical attribute model; Based on the environmental occlusion relationship in the three-dimensional geometric model, calculate the environmental light occlusion factor through a ray tracing algorithm; According to the environmental light occlusion factor and the environmental light intensity, calculate the transparency parameter τ of the AR overlay content, where: τ = k·(Lmax + Lambient) / Lambient, where k is the global adjustment coefficient, Lambient is the environmental light intensity, and Lmax is the maximum brightness of the AR display device; Based on environmental parameters such as temperature and humidity in the physical attribute model, calculate the contrast adjustment coefficient of the AR display device; Generate the AR device adjustment parameters according to the transparency parameter and the contrast adjustment coefficient.
6. The method according to claim 1, wherein The method further includes: Generate a warning message according to the target environmental data and a preset risk assessment rule; Display the warning message on the display screen of the AR device in an AR augmented reality manner.
7. The method according to claim 6, wherein The generating the warning message according to the target environmental data and a preset risk assessment rule includes: Obtain the environmental parameter values in the target environmental data, and compare the environmental parameter values with a preset risk threshold. When the environmental parameter values exceed the risk threshold, determine the risk level; Generate a corresponding warning message according to the risk level.
8. An environment-based perception enhancement system, characterized in that, The system includes: A data acquisition module, configured to acquire N environmental images through a camera of the AR device to obtain an environmental image set, and perform scene classification on the environmental images in the environmental image set to obtain a preliminary scene type; A template matching module, configured to perform optimization processing on the preliminary scene type to obtain a target scene type, and match a corresponding application scene template from a preset scene template library according to the scene type; An environmental parameter determination module, configured to determine the type of environmental parameters to be collected according to the application scene template, and collect environmental data through sensors corresponding to the type of environmental parameters to be collected to obtain target environmental data; An adjustment parameter determination module, configured to generate AR device adjustment parameters according to the target environmental data; An adjustment module, configured to adjust the display parameters of the AR device according to the AR device adjustment parameters.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores multiple instructions, and the instructions are adapted to be loaded and executed by a processor to perform the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory so that the electronic device performs the method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-camera cooperative intelligent early warning method and system for abnormal events
CN120877188A
SLAM method and device based on scene prior and medium
CN121612272A
Display control method and device, head-up display equipment and computer storage medium
CN121657290A