AR live-action navigation method and device and storage medium

By using the preset semantic segmentation model to extract and process the image data collected in real-time in AR real-life navigation, generate a solid mask set, and render a virtual object set, the problem of low navigation accuracy of conventional AR real-life navigation is solved, and real-time recognition and high-precision navigation of dynamic scene objects are realized.

CN119992018AActive Publication Date: 2025-05-13SHENZHEN SMARTCITY TECH DEV GRP CO LTD

Patent Information

Application Number
CN202510451948.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-05-13
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The conventional AR real-life navigation navigation has low accuracy and cannot recognize various objects in real-time dynamic scenes.

Method used

By acquiring the image data collected in real time, using the encoder and decoder of the preset semantic segmentation model to extract and process the image data, obtain a solid mask set, and render the virtual object set in the AR real scene based on this to provide navigation guidance.

Benefits of technology

Real-time recognition of objects in real-time dynamic scenes is achieved, and navigation accuracy and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992018A_ABST
    Figure CN119992018A_ABST
Patent Text Reader

Abstract

The invention discloses an AR live-action navigation method and device and a storage medium, and relates to the technical field of image processing, and the AR live-action navigation method comprises the following steps: obtaining image data collected in real time in an AR live-action navigation process; performing feature extraction on the image data by using a backbone network in an encoder of a preset semantic segmentation model to obtain shallow feature data and deep feature data; performing feature extraction on the deep feature data by using a cavity space convolution pooling pyramid in an encoder to obtain first feature data; performing feature processing on the shallow feature data and the first feature data by using a decoder of a preset semantic segmentation model to obtain an entity mask set corresponding to the image data; rendering based on the entity mask set to obtain a first virtual object set in the AR real scene; navigation guidance is performed based on the rendered and displayed first virtual object set, so that the navigation precision is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an AR real-scene navigation method, device, and storage medium. Background Art

[0002] At present, augmented reality (AR) technology can overlay virtual navigation information on the real road picture, so that users can intuitively understand the route, thereby reducing the difficulty of understanding navigation information. However, conventional AR real-scene navigation simply overlays virtual icons and routes on the real scene, resulting in low navigation accuracy.

[0003] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0004] The main purpose of this application is to provide an AR real-scene navigation method, device and storage medium, aiming to solve the technical problem of low navigation accuracy.

[0005] To achieve the above objectives, the present application proposes an AR real scene navigation method, which includes: Obtain image data collected in real time during AR real-scene navigation; The backbone network in the encoder of the preset semantic segmentation model is used to extract features from the image data to obtain shallow feature data and deep feature data; The deep feature data is extracted using the atrous spatial convolution pooling pyramid in the encoder to obtain the first feature data; Using a decoder of a preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data, to obtain a set of entity masks corresponding to the image data; Based on the entity mask set, a first virtual object set in the AR real scene is rendered; Navigation guidance is provided based on the first set of virtual objects displayed by rendering.

[0006] In one embodiment, the steps of extracting features from image data using a backbone network in an encoder of a preset semantic segmentation model to obtain shallow feature data and deep feature data include: The image data is passed through the first point-by-point convolution module, the first three-by-three convolution module, the attention mechanism module and the second point-by-point convolution module in the backbone network in turn to obtain shallow feature data and deep feature data.

[0007] In one embodiment, the step of extracting features from deep feature data using a dilated spatial convolutional pooling pyramid in an encoder to obtain first feature data includes: The deep feature data are processed respectively by the dilated convolution layer, the third point-by-point convolution module and the first global average pooling layer in the dilated spatial convolution pooling pyramid to obtain the first feature vectors output by the dilated convolution layer, the third point-by-point convolution module and the first global average pooling layer respectively; Concatenate all first eigenvectors in the channel dimension to obtain a first multi-channel eigenvector; The first multi-channel feature vector is convolved through a fourth point-by-point convolution module to obtain first feature data.

[0008] In one embodiment, the step of performing feature processing on the shallow feature data and the first feature data using a decoder of a preset semantic segmentation model to obtain a set of entity masks corresponding to the image data includes: Passing the shallow feature data through the fifth point-by-point convolution module of the decoder of the preset semantic segmentation model to obtain a second feature vector; Processing the first feature data through a first upsampling module to obtain upsampled first feature data; Concatenate the second feature vector with the upsampled first feature data in the channel dimension to obtain a second multi-channel feature vector; Passing the second multi-channel feature vector through the second three-by-three convolution module and the second upsampling module in sequence to obtain a prediction result; Based on the prediction results, a set of entity masks corresponding to the image data is obtained.

[0009] In one embodiment, the step of obtaining a set of entity masks corresponding to the image data based on the prediction result includes: Processing the prediction result using a closed operation filter to obtain a first prediction result; Processing the first prediction result using an open operation filter to obtain a second prediction result; The second prediction result is processed using a Gaussian filter to obtain a post-processing prediction result; Based on the post-processing prediction results, a set of entity masks is determined.

[0010] In one embodiment, based on the entity mask set, the step of rendering to obtain a first set of virtual objects in the AR real scene includes: determining a user's attention threshold relative to an object in the image data; The product of the attention threshold and the area of ​​the image data is used as the area threshold; If it is recognized that the user is paying attention to the first object, select a physical mask with an area greater than or equal to the area threshold from the physical mask set as a first set of objects to be rendered, and render according to relevant information of the first set of objects to be rendered to obtain a first set of virtual objects in the AR real scene; If it is recognized that the user is paying attention to the second object, a physical mask with an area less than the area threshold is selected from the physical mask set as the second set of objects to be rendered, and rendering is performed according to the relevant information of the second set of objects to be rendered to obtain a first set of virtual objects in the AR real scene, wherein the first object is a physical mask with a polygonal area greater than or equal to the area threshold, and the second object is a physical mask with a polygonal area less than the area threshold.

[0011] In one embodiment, before the step of extracting features from image data using a backbone network in an encoder of a preset semantic segmentation model to obtain shallow feature data and deep feature data, the step further includes: Obtain a traffic scene image dataset; Input the traffic scene image dataset into the semantic segmentation model to obtain the prediction results of the semantic segmentation model; Use the loss function to calculate the loss value between the predicted result and the standard preset result of the traffic scene image dataset; If the loss value is greater than or equal to the preset loss value, the weight coefficient of the semantic segmentation model is adjusted according to the loss value to obtain the adjusted semantic segmentation model, and the execution is returned to input the traffic scene image dataset into the semantic segmentation model to obtain the prediction result of the semantic segmentation model; and / or, If the loss value is less than the preset loss value, the preset image segmentation model is output.

[0012] In one embodiment, the step of providing navigation guidance based on the rendered and displayed first virtual object set includes: Get user navigation information; Determine a navigation guidance mark according to the user navigation information and the first virtual object set; Render and display navigation guidance signs on the AR real-scene interface.

[0013] In addition, to achieve the above-mentioned purpose, the present application also proposes an AR real-scene navigation device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the computer program is configured to implement the steps of the AR real-scene navigation method as described above.

[0014] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium. A computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the AR real-scene navigation method as described above are implemented.

[0015] The present application obtains image data collected in real time during AR real-scene navigation; uses the backbone network in the encoder of the preset semantic segmentation model to extract features from the image data to obtain shallow feature data and deep feature data; uses the hollow spatial convolution pooling pyramid in the encoder to extract features from the deep feature data to obtain first feature data; uses the decoder of the preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data to obtain a set of entity masks corresponding to the image data; based on the set of entity masks, renders a first set of virtual objects in the AR real scene; and provides navigation guidance based on the rendered and displayed first set of virtual objects. Since the preset semantic segmentation model is used to extract features from the image data, the objects in the image data are segmented and rendered into the AR real scene, various objects in real dynamic scenes can be recognized in real time, thereby improving the accuracy of navigation. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0018] Figure 1 A flowchart of the first embodiment of the AR real scene navigation method of the present application is provided; Figure 2 This is a schematic diagram of the architecture of the backbone network of the AR real-scene navigation method of this application; Figure 3 A detailed flow chart of step S30 in the first embodiment of the AR real scene navigation method of the present application; Figure 4 A detailed flow chart of step S40 in the first embodiment of the AR real scene navigation method of the present application; Figure 5 A detailed flow chart of step S45 in the first embodiment of the AR real scene navigation method of the present application; Figure 6 A detailed flow chart of step S50 in the first embodiment of the AR real scene navigation method of the present application; Figure 7 A detailed flow chart of step S60 in the first embodiment of the AR real scene navigation method of the present application; Figure 8 A flowchart of the second embodiment of the AR real scene navigation method of the present application is provided; Fig. 9 This is a flowchart of the AR real-scene navigation method for this application; Fig.10 This is a schematic diagram of the device structure of the hardware operating environment involved in the AR real-scene navigation method in the embodiment of the present application.

[0019] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0021] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0022] Traditional 2D map navigation usually requires users to construct a correspondence between the map and the real scene in their minds. This may make it difficult for people with poor sense of direction or unfamiliar with maps to understand and follow navigation instructions. Augmented reality (AR) technology can overlay virtual navigation information on the real road screen, allowing users to intuitively understand the route, reducing the difficulty of reading maps and understanding navigation information. AR is used below to represent augmented reality technology. However, conventional AR real-scene navigation simply overlays virtual icons and routes on the real scene, and cannot recognize various objects in real dynamic scenes in real time, and the navigation accuracy is not high.

[0023] In response to the above problems, the main solution of the present application is: obtaining image data collected in real time during AR real-scene navigation; using the backbone network in the encoder of a preset semantic segmentation model to extract features from the image data to obtain shallow feature data and deep feature data; using the atrous spatial convolution pooling pyramid in the encoder to extract features from the deep feature data to obtain first feature data; using the decoder of the preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data to obtain a set of entity masks corresponding to the image data; based on the set of entity masks, rendering a first set of virtual objects in the AR real scene; and providing navigation guidance based on the rendered and displayed first set of virtual objects.

[0024] The solution of the present application uses a preset semantic segmentation model to extract features from image data, thereby segmenting objects in the image data and rendering them into AR real scenes. It can recognize various objects in real dynamic scenes in real time and improve navigation accuracy.

[0025] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, an AR real scene navigation device, etc. The following takes the AR real scene navigation device as an example to illustrate this embodiment and the following embodiments.

[0026] It should be noted that the first point-by-point convolution module, the second point-by-point convolution module, the third point-by-point convolution module, the fourth point-by-point convolution module and the fifth point-by-point convolution module can be five different convolution modules, or can be five convolution modules with at least one group of modules being the same; the first three-by-three convolution module and the second three-by-three convolution module can be the same or different; the first upsampling module and the second upsampling module can be the same or different.

[0027] Based on this, the present application embodiment provides an AR real scene navigation method, referring to Figure 1 , Figure 1 This is a flowchart of the first embodiment of the AR real-scene navigation method of the present application.

[0028] In this embodiment, the AR real scene navigation method includes steps S10 to S60: Step S10, obtaining image data collected in real time during the AR real-scene navigation process.

[0029] It should be noted that the image data may be traffic scene images collected in real time.

[0030] In this embodiment, the Android platform uses the CameraX library (a set of Jetpack libraries launched by Google to simplify the development of Android device cameras) to collect traffic video streams in real time, and then obtains image data by dividing the traffic video streams into frames, thereby achieving real-time data acquisition.

[0031] Step S20, using the backbone network in the encoder of the preset semantic segmentation model to extract features from the image data to obtain shallow feature data and deep feature data.

[0032] It should be noted that the semantic segmentation model is a model used in the field of computer vision to classify each pixel in an image into a predefined category. For example, in an image of a traffic scene, the semantic segmentation model can mark each pixel as "road", "building", "pedestrian", "vehicle" and other different categories, so that each part of the image has a clear semantic identification. The preset semantic segmentation model is a semantic segmentation model obtained by training the traffic scene image dataset of this embodiment.

[0033] Further, step S20 includes step S21: Step S21, the image data is sequentially passed through the first point-by-point convolution module, the first three-by-three convolution module, the attention mechanism module and the second point-by-point convolution module in the backbone network to obtain shallow feature data and deep feature data.

[0034] It should be noted that Figure 2 This is a schematic diagram of the backbone network architecture.

[0035] The first three-by-three convolution module is a convolution module with a three-by-three convolution kernel, and the attention mechanism module may be an SE (Squeeze-and-Excitation) module, including a global average pooling layer, a fully connected layer, an activation function layer, and a weighted pooling layer.

[0036] Shallow feature data contains detailed information of image data, such as low-level features such as edges and textures, while deep feature data contains semantic information and high-level abstract features of the captured image data.

[0037] In this embodiment, the image data is first passed through a first point-by-point convolution module to change the number of channels to obtain first image data, so that the feature information of different channels is linearly combined to extract more representative features; then a convolution operation is performed on a local area of ​​the first image data through a first three-by-three convolution module to obtain second image data. Compared with point-by-point convolution, the three-by-three convolution kernel has a larger receptive field and can obtain richer local context information, which is helpful to extract more complex visual features.

[0038] Then, the importance weight of each feature channel or spatial position in the second image data is learned through the attention mechanism module, and the features are recalibrated to obtain the third image data. Specifically, the two-dimensional feature map of each channel in the second image data is compressed into a real number through the global average pooling layer in the attention mechanism module, and the real number represents the global information of the corresponding channel feature map, so that the model can capture the dependency between channel feature maps in a global range; then, the features obtained by the global average pooling layer are nonlinearly transformed through a multi-layer perceptron (MLP) composed of two fully connected layers, and the mutual dependency between channels is learned to generate the weight coefficient of each channel; then, the weight coefficient output by the fully connected layer is normalized and compressed to the [0,1] interval to obtain the weight value of each channel, which is used to weight the channels of the second image data. The weight vector composed of the normalized weight coefficients is multiplied by each channel of the second image data to achieve weighting of the feature map channels and obtain the third image data, thereby highlighting the features of important channels and suppressing the features of unimportant channels, so that the model can focus on specific tasks in a biased manner.

[0039] The third image data passes through the second point-by-point convolution module, and the features are again adjusted in channel dimension and feature fusion to obtain the fourth image data. Since the features of the third image data have contained rich local and global information after being processed by the previous modules, the second point-by-point convolution module can further integrate these features and adjust the number of channels according to task requirements to obtain the fourth image data.

[0040] The shallow feature data can be the second image data, or it can be the data obtained by splicing or weighted summing the first image data and the second image data, which can capture the local spatial structure of the image and retain the detail information of the image data; the deep feature data can be the fourth image data, or it can be the data obtained by splicing or weighted summing the first image data, the second image data, the third image data and the fourth image data, and then aggregating the features through a convolution module or a pooling layer, so as to integrate the feature advantages of different levels, so that the deep feature data can more comprehensively capture the semantic information and high-level abstract features of the image data.

[0041] Step S30, using the dilated spatial convolution pooling pyramid in the encoder to extract features from the deep feature data to obtain first feature data.

[0042] It should be noted that Atrous Spatial Pyramid Pooling (ASPP) is an important module for computer vision tasks such as semantic segmentation, which aims to effectively capture multi-scale contextual information. Atrous Spatial Convolution Pooling Pyramid creates multiple branches with different receptive fields by using atrous convolution with different dilation rates, thereby sampling image features at different scales. Based on the standard convolution, atrous convolution inserts holes between the convolution kernel elements, so that the convolution kernel can expand the receptive field without increasing parameters and computation.

[0043] Further, see Figure 3 , step S30 includes steps S31 to S33: Step S31, the deep feature data is processed by the dilated convolution layer, the third point-by-point convolution module and the first global flat pooling layer in the dilated spatial convolution pooling pyramid, respectively, to obtain the first feature vectors output by the dilated convolution layer, the third point-by-point convolution module and the first global flat pooling layer respectively.

[0044] It should be noted that the first global average pooling layer is a pre-set global average pooling layer, which is intended to capture the global information of the feature image.

[0045] In this embodiment, the atrous spatial convolution pooling pyramid includes three three-by-three atrous convolution layers with different atrous rates, and the atrous rates of the three atrous convolution layers are set to 6, 12, and 18, respectively, to cover features of different scale ranges. The receptive field of the atrous convolution layer with a smaller atrous rate is relatively small, which can capture detailed features in the image and is suitable for identifying smaller objects or detailed parts of objects; while the receptive field of the convolution layer with a larger atrous rate is larger, and it focuses more on capturing global contextual information, which helps to understand the overall structure of larger objects or scenes in the image. By using three-by-three atrous convolution layers with atrous rates of 6, 12, and 18 for deep feature data, the first feature vector corresponding to each atrous convolution layer is obtained. At the same time, the first global average pooling layer is used for the deep feature data to average pool the entire feature map in the spatial dimension to obtain a feature vector containing global context information, and then the number of channels of the feature vector is adjusted to the same as the number of channels of the first feature vector output by the dilated convolution layer through point-by-point convolution, and the spatial dimension of the feature vector is adjusted to the same as the spatial dimension of the first feature vector output by the dilated convolution layer through upsampling to obtain the first feature vector output by the first global average pooling layer. In addition, the third point-by-point convolution module is used in parallel to process the deep feature data to obtain the first feature vector output by the third point-by-point convolution module. Thus, multiple first feature vectors are obtained by processing the deep feature data in parallel using three three-by-three dilated convolution layers with dilation rates of 6, 12, and 18, the third point-by-point convolution module, and the first global average pooling layer.

[0046] Step S32: concatenate all first feature vectors in the channel dimension to obtain a first multi-channel feature vector.

[0047] In this embodiment, the first multi-channel feature vector is obtained by splicing the first feature vectors obtained in step S31 in the channel dimension. Thus, the multi-scale local information captured by the dilated convolution layer and the global information provided by the first global average pooling layer complement each other, so that the model can pay attention to the local detail features and grasp the overall scene structure, thereby performing semantic segmentation more accurately.

[0048] Step S33: Perform convolution processing on the first multi-channel feature vector through a fourth point-by-point convolution module to obtain first feature data.

[0049] In this embodiment, the fourth point-by-point convolution module is used to perform feature fusion and dimension adjustment on the first multi-channel feature vector to obtain first feature data. The first feature data integrates multi-scale context information and provides rich and valuable feature representation for subsequent semantic segmentation tasks.

[0050] Step S40, using a decoder of a preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data to obtain a set of entity masks corresponding to the image data.

[0051] It should be noted that the entity mask is the result of segmenting the objects in the image data by a preset semantic segmentation model. Each entity mask corresponds to an object in the image data, and all entity masks are combined into an entity mask set.

[0052] Further, see Figure 4 , step S40 includes steps S41 to S45: Step S41, passing the shallow feature data through the fifth point-by-point convolution module of the decoder of the preset semantic segmentation model to obtain a second feature vector.

[0053] In this embodiment, the fifth point-by-point convolution module is used to perform feature fusion and dimension adjustment on the shallow feature data to obtain the second feature data. This enables the second feature data to have the same number of channels as the subsequent spliced ​​data, thereby ensuring that the two can be smoothly spliced ​​in the channel dimension; and by weighting the features of each channel of the shallow feature data, the feature channels that are more important to the current task are highlighted, making the features after feature fusion more conducive to subsequent upsampling and final prediction; in addition, the overall number of parameters and the amount of calculation of the model can be reduced without seriously affecting the performance of the model.

[0054] Step S42: Process the first feature data through a first upsampling module to obtain upsampled first feature data.

[0055] Step S43: concatenate the second feature vector and the upsampled first feature data in the channel dimension to obtain a second multi-channel feature vector.

[0056] It should be noted that in image processing and deep learning, upsampling is an operation that converts a low-resolution image or feature map to a high-resolution one. Upsampling methods include bilinear interpolation and nearest neighbor interpolation.

[0057] In this embodiment, the first upsampling module uses bilinear interpolation upsampling to amplify the first feature data by four times to obtain upsampled first feature data, and referring to step S32, the upsampled first feature data is concatenated with the second feature vector in the channel dimension to obtain a second multi-channel feature vector. Thus, the preset semantic segmentation model can combine the shallow feature data and the deep feature data for feature extraction, and thus perform semantic segmentation more accurately.

[0058] Step S44, the second multi-channel feature vector is sequentially passed through a second three-by-three convolution module and a second upsampling module to obtain a prediction result.

[0059] It should be noted that the prediction result is a two-dimensional matrix with the same spatial size as the input image. Each element in the matrix corresponds to a pixel of the input image. The value of the matrix element represents the category label to which the pixel belongs. In this embodiment, the input image is image data collected in real time during AR real-scene navigation. For example, in a binary classification task, the value corresponding to the object in the entity mask is 1, and the value corresponding to the background in the entity mask is 0. In a multi-classification task, such as dividing the image into multiple categories such as roads, buildings, vehicles, pedestrians, etc., different categories have different corresponding values ​​in the entity mask.

[0060] Step S45, based on the prediction result, obtaining a set of entity masks corresponding to the image data.

[0061] In this embodiment, the second multi-channel feature vector is convolved through a second three-by-three convolution module to achieve further optimization and adjustment of the spliced ​​second multi-channel feature vector. The receptive field of the model can be fine-tuned without significantly increasing the amount of calculation, so that the model can better handle objects of various scales, enhance the adaptability and robustness of the model to objects of different sizes, and improve the generalization ability of the model in complex scenarios. The second upsampling module is then used to amplify the feature vector after convolution processing by the second three-by-three convolution module by four times using bilinear interpolation upsampling to obtain the prediction result of the image data. Based on the multiple connected pixel areas, prediction categories and confidences corresponding to the prediction categories on the prediction results, the connected pixel areas with confidences less than the preset confidences are filtered out, and the connected pixel areas corresponding to the confidences greater than or equal to the preset confidences are taken as entity masks. The set of multiple entity masks is the entity mask set. The upsampling operation makes the data size of the model output consistent with the image data size of the model input. In addition, the second three-by-three convolution module captures the local features of the second multi-channel feature vector. Upsampling can map the local features to a larger spatial range, which helps to restore the boundaries and details of objects in the image and improve the segmentation accuracy.

[0062] Further, see Figure 5 , step S45 includes steps S451 to S454: Step S451, use a closed operation filter to process the prediction result to obtain a first prediction result.

[0063] Step S452: Use an open operation filter to process the first prediction result to obtain a second prediction result.

[0064] Step S453: Use a Gaussian filter to process the second prediction result to obtain a post-processing prediction result.

[0065] Step S454, determining a set of entity masks based on the post-processing prediction results.

[0066] In this embodiment, when there are complex textures or noises inside the image data, the pixel distribution on the entity mask corresponding to the larger object in the image data will be discontinuous. Therefore, the prediction result is firstly eroded and expanded in sequence by a closed operation filter to fill the small holes in the prediction result, so that the pixel distribution on the prediction result corresponding to the larger object is continuous, and the morphological structure of the pixel distribution on the prediction result corresponding to the object is changed to obtain the first prediction result.

[0067] Then, the opening operation filter is used to perform dilation and erosion operations on the first prediction result in sequence, so as to remove small noise points and burrs in the first prediction result and obtain the second prediction result.

[0068] Finally, a Gaussian filter is applied to smooth the second prediction result to obtain a post-processing prediction result. Each connected pixel region in the post-processing prediction result corresponds to a physical mask of an object, and then a physical mask set is obtained by segmenting each connected pixel region in the post-processing prediction result. In this way, the obtained physical mask can be further smoothed on the basis of the adjustment of morphological operations such as opening operation filters and closing operation filters, and some hard edges that may be generated by the above morphological operations can be blurred, so as to achieve complete segmentation of larger objects in the image data and improve the accuracy of object segmentation in the image data.

[0069] Step S50: Rendering a first virtual object set in the AR real scene based on the entity mask set.

[0070] It should be noted that each pixel value in the physical mask represents the category to which the physical mask belongs. For example, if you want to identify people and vehicles in the scene, the pixel area of ​​the physical mask corresponding to the people and vehicles has its specific value. The first virtual object set is an AR real scene display of the object set corresponding to the physical mask set.

[0071] In this embodiment, in the recognition scene of people and vehicles, the AR development framework is used to initialize the AR real scene, obtain the image data of the device camera and the posture information of the device, such as position and direction, to determine the position and display mode of the people and vehicles in the real scene. The pixel coordinates of the physical mask corresponding to the people and vehicles are mapped to the three-dimensional space coordinates of the AR real scene. According to the category represented by the pixel area in the physical mask, the corresponding virtual object is created, and the corresponding virtual object is created in the AR real scene for the physical mask of the person, and the corresponding virtual object is created in the AR real scene for the physical mask of the vehicle. Then, using the rendering function provided by the AR development framework, the virtual objects corresponding to the people and vehicles are rendered to the corresponding three-dimensional space coordinates in the AR real scene, and the first virtual object set, that is, the virtual object set of people and vehicles is obtained, thereby realizing the AR display of objects in the real image data.

[0072] In another implementation of this embodiment, by receiving selection information of a user for an object of a specific category, a physical mask of a corresponding category that meets the selection information is selected from a physical mask set to render a virtual object on an AR real scene.

[0073] Further, see Figure 6 , step S50 includes steps S51 to S54: Step S51, confirming the user's attention threshold relative to the object in the image data.

[0074] It should be noted that the attention threshold is a value greater than 0 and less than 1 set by the user, which can screen out objects that meet the user's expectations for AR real scene rendering.

[0075] In step S52, the product of the attention threshold and the area of ​​the image data is used as the area threshold.

[0076] Step S53: If it is recognized that the user is paying attention to the first object, select a physical mask with an area greater than or equal to the area threshold from the physical mask set as the first set of objects to be rendered, and render according to the relevant information of the first set of objects to be rendered to obtain a first set of virtual objects in the AR real scene.

[0077] Step S54: If it is recognized that the user is paying attention to the second object, a physical mask with an area less than the area threshold is selected from the physical mask set as the second set of objects to be rendered, and rendering is performed according to the relevant information of the second set of objects to be rendered to obtain a first set of virtual objects in the AR real scene, wherein the first object is a physical mask with a polygonal area greater than or equal to the area threshold, and the second object is a physical mask with a polygonal area less than the area threshold.

[0078] It should be noted that the data in the first virtual object set is not less than one.

[0079] In this embodiment, there is a car and a person in the image data, and an attention threshold is received through the user terminal setting interface. The area threshold is obtained by calculating the product of the attention threshold and the area of ​​the image data. The polygonal area of ​​the pixel area of ​​the physical mask corresponding to the vehicle is greater than the area threshold, and it belongs to the first object. The polygonal area of ​​the pixel area of ​​the physical mask corresponding to the person is less than the area threshold, and it belongs to the second object. When the user terminal setting interface receives that the user is paying attention to the first object, refer to the method in step S50, use the pixel area of ​​the physical mask corresponding to the vehicle to complete the rendering of the corresponding virtual object in the AR real scene, and do not render the physical mask corresponding to the person on the AR real scene, so as to obtain a first virtual object set, so as to achieve selective attention to objects with physical masks and meet the user's demand selection.

[0080] In another implementation of the present embodiment, there is a car and a person in the image data, and an attention threshold is received through the user terminal setting interface. The area threshold is obtained by calculating the product of the attention threshold and the area of ​​the image data. The polygonal area of ​​the pixel area of ​​the physical mask corresponding to the vehicle is greater than the area threshold, and it belongs to the first object. The polygonal area of ​​the pixel area of ​​the physical mask corresponding to the person is less than the area threshold, and it belongs to the second object. When the user terminal setting interface receives that the user is paying attention to the second object, refer to the method in step S50, use the pixel area of ​​the physical mask corresponding to the person to complete the rendering of the corresponding virtual object in the AR real scene, and do not render the physical mask corresponding to the vehicle on the AR real scene, so as to obtain a first virtual object set, so as to achieve selective attention to objects with physical masks and meet the user's demand selection.

[0081] Step S60: providing navigation guidance based on the rendered and displayed first virtual object set.

[0082] Further, see Figure 7 , step S60 includes steps S61-S62: Step S61, obtaining user navigation information.

[0083] It should be noted that the user navigation information is the navigation route information selected by the user, including the starting point, the end point, and the route track from the starting point to the end point.

[0084] Step S62: determine a navigation guide mark according to the user navigation information and the first virtual object set, and render and display the navigation guide mark on the AR real scene interface.

[0085] It should be noted that the navigation guidance mark is a virtual mark information displayed on the AR real scene to provide navigation guidance information to the user, such as arrows, signs, positioning points, etc.

[0086] In this embodiment, one or more virtual objects are selected from the first virtual object set as navigation targets, and the map information in the AR real scene or the scene map constructed in real time by the simultaneous localization and mapping (SLAM) technology is used to receive user navigation information through the user front-end interface, so as to plan the path from the user's current position to the navigation target, and then present navigation instructions to the user through the AR interface, such as displaying guiding signs such as arrows and signs on the virtual objects, or drawing the navigation path in the user's perspective. At the same time, voice prompts can also be combined to inform the user of information such as the direction and distance of travel, so as to help the user reach the navigation target smoothly and improve the accuracy of navigation.

[0087] This embodiment is implemented based on the first embodiment. In this embodiment, before step S20 of the AR real scene navigation method, refer to Figure 8 , further comprising steps B10 to B50: Step B10, obtaining a traffic scene image dataset.

[0088] In this embodiment, traffic scene image data with polygon annotation is obtained, and a traffic scene image data set is produced, which includes information such as the original image, mask, and category of the image, wherein the shape of the mask corresponds to the polygon shape of the polygon annotation. In addition, before polygon annotation, data enhancement can be performed on the image in the traffic scene image data, including changing brightness, adding noise, adding random points, translation, flipping, etc., so as to make the image more random, so as to improve the generalization ability of the model. At the same time, the model hyperparameters are determined through the configuration table.

[0089] Step B20, inputting the traffic scene image dataset into the semantic segmentation model to obtain the prediction result of the semantic segmentation model.

[0090] Step B30, using a loss function to calculate a loss value between the prediction result and a standard preset result of the traffic scene image dataset.

[0091] It should be noted that the standard preset result is the data obtained after the traffic scene image dataset is converted into a data format corresponding to the prediction result, which is used as the judgment standard of the prediction result.

[0092] The loss function includes at least one of a cross entropy loss, a regularization loss, and a Dice loss (a type of loss function).

[0093] Step B40, if the loss value is greater than or equal to the preset loss value, adjust the weight coefficient of the semantic segmentation model according to the loss value to obtain the adjusted semantic segmentation model, and return to execute the input of the traffic scene image dataset into the semantic segmentation model to obtain the prediction result of the semantic segmentation model.

[0094] Step B50, and / or, if the loss value is less than a preset loss value, output a preset image segmentation model.

[0095] In this embodiment, the loss function includes cross entropy loss, regularization loss, and Dice loss. The traffic scene image data set is input into the semantic segmentation model to obtain the prediction result output by the semantic segmentation model. The data in the traffic scene image data set is formatted according to the data format of the prediction result to obtain the standard preset result. The loss function is then used to calculate the loss value between the prediction result and the standard preset result. If the loss value is greater than or equal to the preset loss value, the optimizer is instructed to adjust the weight coefficient of the semantic segmentation model according to the loss value calculated by the model. The above steps are repeated until the loss value is less than the preset loss value, the model training is completed, and the preset semantic segmentation model is obtained, thereby providing a model basis for achieving high-precision image object segmentation.

[0096] In another implementation of the present embodiment, a traffic scene image dataset is input into a semantic segmentation model to obtain a prediction result output by the semantic segmentation model. The data in the traffic scene image dataset is formatted according to the data format of the prediction result to obtain a standard preset result. A loss function is then used to calculate the loss value between the prediction result and the standard preset result. When the loss value is compared and is less than the preset loss value, model training is completed to obtain a preset semantic segmentation model, thereby providing a model basis for achieving high-precision image object segmentation.

[0097] For example, to help understand the implementation process of the AR real scene navigation method obtained by combining this embodiment with the above-mentioned embodiment 1, please refer to Fig. 9 , Fig. 9 A brief flowchart of an AR real-scene navigation method is provided, specifically: The backbone network of the encoder receives input image data, performs feature extraction on the image data, obtains deep feature data and shallow feature data, applies the third point-by-point convolution module, the first global average pooling layer and the three hole convolution layers to process the deep feature data respectively, obtains multiple first feature vectors, splices these first feature vectors in the channel dimension to obtain a first multi-channel feature vector, and convolves the first multi-channel feature vector through the fourth point-by-point convolution module to obtain the first feature data; the decoder receives the shallow feature data from the backbone network, performs feature extraction on the shallow feature data through the fifth point-by-point convolution module to obtain a second feature vector, splices the second feature vector with the first multi-channel feature vector upsampled by the first upsampling module in the channel dimension to obtain a second multi-channel feature vector, and processes the second multi-channel feature vector in sequence using the second three-by-three convolution module and the second upsampling module to obtain the prediction result of the preset semantic segmentation model.

[0098] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the AR real-scene navigation method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0099] The present application provides an AR real-scene navigation device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the AR real-scene navigation method in the above-mentioned embodiment one.

[0100] Reference below Fig.10 , which shows a schematic diagram of the structure of an AR real-scene navigation device suitable for implementing an embodiment of the present application. The AR real-scene navigation device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDA, Personal Digital Assistant), tablet computers (PAD, Portable Application Description), portable multimedia players (PMP, Portable Media Player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Fig.10 The AR real-scene navigation device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0101] like Fig.10As shown, the AR real scene navigation device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage device 1003 to the random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the AR real scene navigation device are also stored. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the AR real-scene navigation device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows an AR real-scene navigation device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or provided alternatively.

[0102] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0103] The AR real-scene navigation device provided by the present application adopts the AR real-scene navigation method in the above embodiment, which can solve the technical problem of low navigation accuracy. Compared with the prior art, the beneficial effects of the AR real-scene navigation device provided by the present application are the same as the beneficial effects of the AR real-scene navigation method provided by the above embodiment, and the other technical features in the AR real-scene navigation device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0104] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0105] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0106] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the AR real-scene navigation method in the above-mentioned embodiment.

[0107] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM, Erasable Programmable ReadOnly Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequencies (RF, Radio Frequency), etc., or any suitable combination of the above.

[0108] The computer-readable storage medium may be included in the AR real-scene navigation device; or may exist independently without being assembled into the AR real-scene navigation device.

[0109] The computer-readable storage medium carries one or more programs. When the one or more programs are executed by the AR real-scene navigation device, the AR real-scene navigation device: obtains image data collected in real time during the AR real-scene navigation process; uses the backbone network in the encoder of the preset semantic segmentation model to extract features from the image data to obtain shallow feature data and deep feature data; uses the hole space convolution pooling pyramid in the encoder to extract features from the deep feature data to obtain first feature data; uses the decoder of the preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data to obtain a set of entity masks corresponding to the image data; based on the set of entity masks, renders a first set of virtual objects in the AR real scene; and provides navigation guidance based on the rendered and displayed first set of virtual objects.

[0110] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages, such as Java, Smalltalk, C++, and conventional procedural programming languages, such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0111] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0112] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0113] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned AR real-scene navigation method, and can solve the technical problem of low navigation accuracy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the AR real-scene navigation method provided in the above-mentioned embodiment, and will not be repeated here.

[0114] The above are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.

Claims

1. An AR real scene navigation method, characterized in that: The AR real scene navigation method comprises: Obtain image data collected in real time during AR real-scene navigation; Using a backbone network in an encoder of a preset semantic segmentation model to extract features from the image data, to obtain shallow feature data and deep feature data; Extracting features from the deep feature data using a dilated spatial convolutional pooling pyramid in the encoder to obtain first feature data; Using the decoder of the preset semantic segmentation model to perform feature processing on the shallow feature data and the first feature data to obtain a set of entity masks corresponding to the image data; Based on the entity mask set, rendering obtains a first virtual object set in the AR real scene; Navigation guidance is provided based on the first set of virtual objects displayed by rendering.

2. The AR real scene navigation method according to claim 1, characterized in that: The step of extracting features from the image data using the backbone network in the encoder of the preset semantic segmentation model to obtain shallow feature data and deep feature data comprises: The image data is sequentially passed through the first point-by-point convolution module, the first three-by-three convolution module, the attention mechanism module and the second point-by-point convolution module in the backbone network to obtain the shallow feature data and the deep feature data.

3. The AR real scene navigation method according to claim 2, characterized in that: The step of extracting features from the deep feature data using the atrous spatial convolutional pooling pyramid in the encoder to obtain first feature data comprises: The deep feature data are processed respectively by the dilated convolution layer, the third point-by-point convolution module and the first global flat pooling layer in the dilated spatial convolution pooling pyramid to obtain first feature vectors output by the dilated convolution layer, the third point-by-point convolution module and the first global flat pooling layer; Concatenating all the first feature vectors in the channel dimension to obtain a first multi-channel feature vector; The first multi-channel feature vector is convolved through a fourth point-by-point convolution module to obtain the first feature data.

4. The AR real scene navigation method according to claim 1 or 3, characterized in that: The step of performing feature processing on the shallow feature data and the first feature data by using the decoder of the preset semantic segmentation model to obtain a set of entity masks corresponding to the image data comprises: Passing the shallow feature data through the fifth point-by-point convolution module of the decoder of the preset semantic segmentation model to obtain a second feature vector; Processing the first feature data through a first upsampling module to obtain upsampled first feature data; Concatenate the second feature vector and the upsampled first feature data in the channel dimension to obtain a second multi-channel feature vector; Passing the second multi-channel feature vector sequentially through a second three-by-three convolution module and a second upsampling module to obtain a prediction result; Based on the prediction result, a set of entity masks corresponding to the image data is obtained.

5. The AR real scene navigation method according to claim 4, characterized in that: The step of obtaining a set of entity masks corresponding to the image data based on the prediction result comprises: Processing the prediction result using a closed operation filter to obtain a first prediction result; Processing the first prediction result using an open operation filter to obtain a second prediction result; Processing the second prediction result using a Gaussian filter to obtain a post-processing prediction result; Based on the post-processing prediction result, the entity mask set is determined.

6. The AR real scene navigation method according to claim 1, characterized in that: The step of rendering a first virtual object set in an AR real scene based on the entity mask set includes: determining a user's attention threshold relative to an object in the image data; The product of the attention threshold and the area of ​​the image data is used as the area threshold; If it is recognized that the user is paying attention to the first object, select a physical mask with an area greater than or equal to the area threshold from the physical mask set as a first set of objects to be rendered, and render according to relevant information of the first set of objects to be rendered to obtain the first set of virtual objects in the AR real scene; If it is recognized that the user is paying attention to the second object, a physical mask with an area smaller than the area threshold is selected from the physical mask set as the second set of objects to be rendered, and rendering is performed according to the relevant information of the second set of objects to be rendered to obtain the first set of virtual objects in the AR real scene, wherein the first object is a physical mask with a polygonal area greater than or equal to the area threshold, and the second object is a physical mask with a polygonal area smaller than the area threshold.

7. The AR real scene navigation method according to claim 1, characterized in that: Before the step of extracting features from the image data using the backbone network in the encoder of the preset semantic segmentation model to obtain shallow feature data and deep feature data, the method further includes: Obtain a traffic scene image dataset; Inputting the traffic scene image dataset into a semantic segmentation model to obtain a prediction result of the semantic segmentation model; Calculate the loss value between the prediction result and the standard preset result of the traffic scene image dataset using a loss function; If the loss value is greater than or equal to a preset loss value, adjusting the weight coefficient of the semantic segmentation model according to the loss value to obtain an adjusted semantic segmentation model, and returning to execute the step of inputting the traffic scene image dataset into the semantic segmentation model to obtain a prediction result of the semantic segmentation model; and / or, If the loss value is less than the preset loss value, the preset image segmentation model is output.

8. The AR real scene navigation method according to claim 1, characterized in that: The step of providing navigation guidance based on the first set of virtual objects displayed by rendering comprises: Get user navigation information; Determine a navigation guidance mark according to the user navigation information and the first set of virtual objects; The navigation guidance mark is rendered and displayed on the AR real scene interface.

9. An AR real-scene navigation device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the AR real scene navigation method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the AR real scene navigation method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • High-precision semantic segmentation method for automatic driving road scene

    CN117649526A

  • Blind person intelligent navigation system and method based on image semantic segmentation

    CN118298170A

  • Image processing apparatus, image encoder, image printer, and method for them

    JP2003158635A

  • Target detection method and apparatus

    WO2021254205A1

Cited By

  • Complex farmland visual navigation method and system based on lightweight segmentation residual network

    CN120869146A