Method and system for publishing advertisement in virtual reality scene based on mobile live-action cognition
By combining deep learning and SLAM technologies, a high degree of integration and personalized adaptation between virtual advertising and real-world scenes has been achieved, solving the problems of integration and dynamic optimization in virtual reality advertising systems and improving advertising effectiveness and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-03-10
AI Technical Summary
Existing virtual reality advertising systems lack deep real-world perception capabilities, resulting in poor integration of virtual advertisements with the real environment, inability to personalize and adapt them, and a lack of dynamic optimization mechanisms, leading to insufficient accuracy and effectiveness evaluation of advertising.
We employ deep learning models for semantic segmentation and object recognition of real-world data, combine SLAM technology to achieve stable positioning and lighting matching of virtual advertisements, dynamically generate personalized advertising content through real-time user data analysis, and establish a reinforcement learning model to optimize advertising strategies.
It achieves a high degree of integration between virtual advertising and real-world scenarios, providing a personalized advertising experience, improving advertising relevance and user engagement, and enhancing conversion rates and ROI through dynamic optimization.
Smart Images

Figure CN121639276A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of virtual reality and digital advertising technology, and in particular to a method and system for advertising in virtual reality scenarios based on mobile real-world cognition. Background Technology
[0002] Currently, some solutions have attempted to overlay virtual advertising content onto the real environment. For example, they capture real-world scenes using mobile device cameras and display preset advertising information on them. Some systems also use basic image recognition technology to detect specific objects in the scene, thereby triggering the corresponding advertising display. These technologies are mostly based on fixed markers or simple scene features for advertising, which to some extent achieves the integration of virtual content with the real environment.
[0003] Most systems lack a deep understanding of the real-world environment, capable only of simple object recognition. They fail to comprehend key elements such as semantic information, spatial structure, and ambient lighting, resulting in poor integration of virtual ads with the real environment and often making the ad content appear stiff and unnatural. Furthermore, existing technologies struggle to personalize ads based on user attributes and real-time behavior, limiting the accuracy of ad delivery. In addition, traditional systems lack effective ad performance evaluation and dynamic optimization mechanisms, failing to adjust ad strategies in real time based on user feedback. Regarding mobile management, existing solutions offer limited functionality, making it difficult for advertisers to achieve refined delivery control and performance monitoring.
[0004] Therefore, in response to the problems mentioned above, this invention proposes a method and system for advertising in virtual reality scenarios based on mobile real-world cognition. Summary of the Invention
[0005] To overcome the problems of insufficient integration between virtual reality advertising and the real-world environment, low personalization, and lack of dynamic optimization mechanisms, this invention proposes a method and system for advertising in virtual reality scenes based on mobile real-world cognition. Through deep scene understanding and intelligent adaptation technology, it achieves a high degree of integration between advertising content and the real-world environment, and provides personalized advertising experience and dynamic optimization functions.
[0006] The technical solution of this invention is: a method for advertising in a virtual reality scene based on mobile real-world cognition, comprising the following steps: S1 collects real-time real-scene data of the user's surrounding environment through the mobile device's camera and sensors. The real-scene data includes scene images, location information (such as GPS coordinates), spatial structure (such as depth map), and lighting conditions (such as brightness and color temperature). Deep learning models are used to perform semantic segmentation and object recognition on the real-world data, thereby identifying key elements in the scene, including human bodies, clothing stores, fitting rooms, and public areas, and extracting scene features (such as color histograms and texture features). The use of deep learning models includes image analysis using convolutional neural networks, such as using a pre-trained YOLO model to achieve real-time segmentation and detection, and combining sensor data fusion (such as fusing gyroscope and accelerometer data) to improve recognition accuracy. At the same time, based on the recognition results of real-scene data, user attribute data is obtained, including gender, age and body shape data (such as estimating height and weight through image analysis), as well as real-time behavioral data, including gaze direction (through eye tracking) and gestures (captured through a camera). S2, based on the recognition results of real-scene data and user attribute data, dynamically select or generate clothing advertising content from the advertising database. The clothing advertising content includes 3D clothing models, video advertisements, interactive try-on interfaces, and promotional information. Among them, dynamically selecting or generating clothing advertising content from the advertising database includes: using recommendation algorithms to prioritize displaying clothing styles that match the user's body type and preferences based on user attributes and real-time behavioral data, such as recommending a suitable size based on the user's body type, or recommending similar styles based on historical behavior; Virtual reality scenes are built on mobile devices, and clothing advertising content is seamlessly overlaid into the real scene using augmented reality (AR) or mixed reality (MR) technology. The construction of virtual reality scenes includes adjusting the material and shadows of virtual clothing according to the real scene lighting, and using SLAM technology to achieve stable positioning of virtual objects in the real scene. Preferably, the construction of the virtual reality scene includes: dynamically adjusting the display angle and size of the virtual advertisement according to the user's position to optimize the viewing experience, such as automatically scaling the advertisement model when the user moves to ensure that the advertisement is always in the best viewing angle; S3, In a virtual reality scene, the clothing advertisement content is published, which associates the advertisement with a specific object in the real scene (such as attaching an advertisement label to a clothing rack), and is displayed in real time through a mobile device display screen, allowing users to interact with the clothing advertisement content through gestures, voice or touch screen, including rotating the clothing model, virtual try-on, and obtaining purchase links; Among them, allowing users to interact with clothing advertising content includes: enabling virtual try-on through gesture recognition, wherein the mobile device camera captures user gestures (such as opening a palm or swiping), and renders clothing models onto the user's body in real time in a virtual reality scene; S4 collects user interaction data (such as click-through rate, dwell time, and number of try-ons) and scene data (such as environmental complexity) in real time to evaluate advertising effectiveness; The evaluation of advertising effectiveness includes: calculating click-through rate, dwell time, and number of try-ons; and using reinforcement learning models to dynamically adjust advertising strategies in conjunction with scene complexity data, such as adjusting the frequency of ad display based on user attention allocation; and using machine learning algorithms to dynamically optimize ad content, location, and display timing to improve conversion rate. Preferably, the optimization cycle is set to real-time or near real-time (e.g., updated every 5 seconds). S5 provides advertising management functions through the management interface on the mobile user terminal, including customizing advertising content (such as uploading 3D clothing models and setting up promotional activities), specifying target scenarios (such as only publishing in the fitting room area), setting budget and delivery rules (such as based on user gender or time period), real-time monitoring of advertising performance, and generating analysis reports. The mobile user interface is provided through a graphical user interface, allowing advertisers to upload 3D clothing models, set up promotional activities, and remotely update advertising content based on real-time data. It also supports multi-platform integration (such as connecting to e-commerce systems via API).
[0007] This invention proposes a system for advertising in virtual reality scenarios based on mobile reality cognition, comprising: The mobile user terminal includes a real-scene cognition module, a virtual reality rendering module, an advertising interaction module, and a publishing management interface. The mobile user terminal is used to collect real-scene data, generate virtual reality scenes, display clothing advertising content, and provide user interaction and management functions. The server side includes an advertising database, a scene analysis engine, an advertising scheduling engine, and a user management module. The server side is used to store advertising content, analyze real-world data, schedule advertising releases, and manage user data. The network communication module is used to realize data exchange between the mobile user terminal and the server, and supports low-latency transmission; The real-world perception module uses a deep learning model to perform semantic segmentation and object recognition on real-world data; the virtual reality rendering module uses a graphics rendering engine to overlay clothing advertising content onto the real-world scene; and the advertising scheduling engine selects advertising content based on the real-world perception results and user attributes. Preferably, the mobile user terminal is a smartphone or tablet computer equipped with a camera, GPS sensor, gyroscope and accelerometer, and the real-scene cognition module executes a deep learning model through a built-in processor to identify scene elements in real time.
[0008] Preferably, the server-side advertising database stores 3D clothing models and metadata, the scene analysis engine uses machine learning algorithms to perform advanced scene analysis and return advertising recommendation results, and the advertising scheduling engine integrates a rule engine to optimize advertising delivery based on the target scenes and delivery rules set by the advertisers.
[0009] The beneficial effects of this invention are: 1. This invention employs semantic segmentation and object recognition technology based on convolutional neural networks, which can accurately identify key elements such as human bodies and clothing stores in real-world scenes. Combined with SLAM technology, it achieves stable positioning of virtual clothing advertisements in real-world scenes. Furthermore, by adjusting the material and shadows of virtual clothing through real-time lighting matching technology, it ensures that the virtual advertisements are visually highly consistent with the real environment. At the same time, it utilizes multimodal data fusion to enhance environmental understanding capabilities, significantly improving the naturalness and immersiveness of the integration between the advertising content and the real-world scene, making the virtual clothing advertisements appear as if they truly exist in the environment.
[0010] 2. This invention, by collecting and analyzing user attribute data (such as gender, age, and body type) and behavioral data (such as gaze direction and gestures) in real time, combined with deep learning recommendation algorithms, can dynamically generate clothing advertising content that highly matches the user's body shape characteristics and style preferences. It supports virtual try-on functionality based on skeleton tracking technology, enabling advertising content to fit the user's body contours in real time. At the same time, it allows advertisers to accurately set target scenarios and delivery rules through the management interface, achieving comprehensive personalized adaptation from advertising content to delivery scenarios, and significantly improving the relevance of advertisements and user engagement.
[0011] 3. This invention establishes a comprehensive dynamic optimization mechanism. By collecting user interaction data (click-through rate, dwell time, number of try-ons) and scene complexity data in real time, it uses a reinforcement learning model to dynamically adjust advertising strategies. This supports near real-time (e.g., every 5 seconds) optimization of advertising content, display position, and timing. Simultaneously, by combining advertising performance evaluation data, it continuously optimizes the delivery strategy through machine learning algorithms, thereby improving advertising conversion rate and return on investment, and solving the problem of traditional virtual advertising lacking continuous optimization capabilities. Attached Figure Description
[0012] Figure 1 The diagram shown is a schematic representation of the system framework of this invention. Figure 2 The diagram shown illustrates the workflow of this invention. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of the present invention, but not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0014] Please see Figure 1This invention provides an embodiment of a system for advertising in a virtual reality scene based on mobile reality cognition, comprising: In this embodiment, the mobile user terminal is a high-performance smartphone equipped with a 12-megapixel camera, GPS module, gyroscope and accelerometer sensor. The real-scene cognition module is integrated into the mobile application and uses a pre-trained convolutional neural network model (ResNet-50 and YOLOv4 architecture) for semantic segmentation and object recognition. ResNet-50 is used to extract scene features (such as color histograms and texture features), while YOLOv4 is used to detect objects such as humans, clothing stores, and fitting rooms in real time. The recognition accuracy is improved to over 90% through sensor data fusion (e.g., fusing gyroscope data to correct image offset). At the same time, this module captures user gestures (such as open palms) and eye-tracking data (using the device's front-facing camera to analyze the direction of gaze) through the camera, and combines image analysis to estimate user body shape data (such as height and weight). All processing is performed locally on the device to reduce latency, and outputs structured data (such as bounding box coordinates and feature vectors) for use by the virtual reality rendering module.
[0015] The virtual reality rendering module is built on the Unity engine and ARKit / ARCore framework. It uses SLAM technology to achieve stable positioning of virtual objects in the real scene. It simulates the material and shadow of virtual clothing through the Phong lighting model to match the lighting conditions of the real scene (such as brightness and color temperature). For example, it adjusts the reflectivity of clothing under indoor lighting. At the same time, it dynamically adjusts the display angle and size of advertisements according to the user's position (such as automatically scaling the 3D model when the user moves). The rendering frame rate is kept above 30fps to ensure a smooth experience.
[0016] The advertising interaction module integrates a gesture recognition library (such as skeleton tracking technology) and a voice recognition engine, allowing users to interact with advertisements through gestures (such as sliding and rotating the clothing model) or voice commands (such as "try on this dress"). The virtual try-on function achieves high-precision fitting by rendering a 3D clothing model onto the user's body skeleton points in real time, with an error of less than 2 centimeters. At the same time, interaction data (such as click events) is captured and cached locally through touch screen controls.
[0017] The publishing management interface is provided in the form of a graphical user interface, allowing advertisers to upload 3D clothing models, set delivery rules (such as targeting female users or publishing in fitting room areas), monitor real-time metrics (such as click-through rate), and synchronize data with the server via API. The interface is designed to be responsive to different mobile device screens.
[0018] In this embodiment, the server is deployed on a cloud platform, and the advertising database uses a hybrid storage of MySQL and MongoDB. MySQL stores structured data (such as user attributes and advertising metadata), while MongoDB stores unstructured data (such as 3D clothing models and video files). Database index optimization supports millisecond-level queries, such as quickly retrieving matching clothing styles based on user body shape data.
[0019] The scene analysis engine uses deep learning models for advanced scene analysis, including semantic segmentation (such as scene recognition complexity calculated by image entropy) and behavior prediction (such as inferring user preferences based on historical data). The analysis results are returned to the mobile user terminal.
[0020] The ad scheduling engine integrates recommendation algorithms to dynamically select ad content based on user attributes (such as age and body type) and real-time behavioral data (such as gaze direction). For example, it can recommend popular clothing by calculating user similarity through collaborative filtering, or generate personalized lists through deep neural network models. At the same time, the scheduling engine has a built-in rule engine to execute the ad placement rules set by advertisers (such as budget constraints) and optimize the timing of ad display.
[0021] The user management module handles authentication and data privacy, encrypts and stores user data, and provides APIs for integration with third-party systems (such as e-commerce platforms) to ensure secure data exchange.
[0022] In this embodiment, the network communication module uses 5G and Wi-Fi 6 protocols to realize data exchange between the mobile user terminal and the server terminal, and supports low-latency transmission.
[0023] Please see Figure 2 This invention provides an embodiment of a method for advertising in a virtual reality scene based on mobile reality cognition: (1) The mobile user terminal collects real-world data by capturing scene images, location information (GPS coordinates), spatial structure (depth map) and lighting conditions (brightness and color temperature) in real time through cameras and sensors. Then, the real-world cognition module uses convolutional neural network models (ResNet-50 and YOLOv4) to perform semantic segmentation and object recognition, identify key elements such as human body and clothing store and extract scene features (such as color histogram), and at the same time combine sensor data fusion to improve accuracy, and obtain user attribute data (gender, age, body type) and real-time behavior data (eye direction and gestures), and output the cognition results to the server.
[0024] (2) After receiving the data, the scene analysis engine on the server side performs advanced analysis (such as quantifying the environmental complexity through image entropy). The advertising scheduling engine, based on the cognitive results and user attributes, uses recommendation algorithms to dynamically select or generate clothing advertising content (such as 3D clothing models or video ads) from the advertising database, and prioritizes displaying styles that match the user's body type and preferences. For example, it recommends a suitable size based on the user's body type. The scheduling results are returned to the mobile user terminal.
[0025] (3) The virtual reality rendering module of the mobile user terminal constructs a virtual reality scene based on the returned data, overlays the advertising content onto the real scene using augmented reality (AR) or mixed reality (MR) technology, uses SLAM technology to achieve stable positioning of virtual objects, and adjusts the virtual clothing material and shadows according to the real scene lighting through the Phong lighting model. At the same time, it dynamically adjusts the display angle and size to optimize the viewing experience, and then publishes the advertisement in the virtual scene, so that the advertisement is associated with specific objects in the real scene (such as clothing racks) and displayed in real time through the display screen.
[0026] (4) Users interact with advertising content through the advertising interaction module, such as virtual try-on through gesture recognition, or purchase links through voice commands. Interaction data (click rate, dwell time, number of try-ons) are collected in real time and uploaded to the server.
[0027] (5) The server-side advertising scheduling engine uses a reinforcement learning model to evaluate the advertising effect and dynamically optimizes the advertising content, location and display timing by combining scene data (such as environmental complexity). For example, the display frequency is adjusted according to user attention and the optimization cycle is set to near real-time (updated every 5 seconds).
[0028] (6) Finally, the mobile client's publishing management interface allows advertisers to manage advertising campaigns, including uploading content, setting rules, monitoring performance and generating reports, and synchronizing with the server through the network module to complete the entire publishing cycle.
[0029] This invention provides a comparative example: This comparative example simulates real-world apparel retail scenarios through experimental environments, including both indoor shopping malls and outdoor plazas, with lighting conditions varying from 100 lux (dim) to 1000 lux (bright).
[0030] The experimental equipment used a standard mobile client configuration. The experimental data included 1000 users (aged 18-60, gender ratio 1:1, body type distribution based on BMI classification), and the advertising database contained 500 clothing items (3D models and video ads). The experiment lasted 4 weeks. Metrics included click-through rate (CTR), user dwell time (seconds), number of try-ons, conversion rate (percentage of purchase intention), system latency (milliseconds), and integration score (1-5 points, assessed by experts for consistency between virtual ads and real-world visuals).
[0031] in: Example 1 uses the present invention.
[0032] Comparative Example 1 uses an existing technology system with a basic AR advertising framework. It only uses simple image tag recognition, without deep learning cognition. The advertising content is statically preset, without personalized recommendations or dynamic optimization, and the interaction is limited to basic touch clicks.
[0033] Comparative Example 2 improves upon existing technology by adding simple CNN recognition, but lacks sensor data fusion or SLAM localization. Ad scheduling is based on a rule engine and has limited interactive functions (no virtual try-on).
[0034] This experiment tested the overall performance of the invention under standard illumination (500 lux).
[0035]
[0036] As shown in the table above, under standard lighting conditions, this invention outperforms the two comparative examples in all indicators. The virtual content displayed in Comparative Examples 1 and 2 is awkwardly and uncoordinated with the real-world scene. This experiment demonstrates that under ideal lighting conditions, the real-world perception and virtual reality fusion technology of this invention can significantly improve the effectiveness of clothing advertising and the user experience.
[0037] This experiment tested the overall performance of the invention in a real-world scenario with drastic fluctuations in light intensity (randomly varying from 200 to 1000 lux, simulating the transition from an indoor shaded area to a brightly lit area).
[0038]
[0039] As shown in the table above, the performance of all systems decreased under high light variation conditions. However, the decrease in performance of the present invention was significantly smaller than that of the comparative examples. Specifically, the fusion score of Example 1 decreased by only 0.2 points compared to the standard lighting experiment, while Comparative Examples 1 and 2 decreased by 0.5 points and 0.5 points respectively. This demonstrates that the Phong lighting model and real-time lighting adaptation mechanism of the present invention effectively mitigated the impact of lighting changes on visual quality. The data in the table proves that even under high light variation conditions, the personalized recommendation and interactive functions of the present invention can still effectively promote user participation.
[0040] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for advertisement publishing in a virtual reality scene based on mobile real scene cognition, characterized in that, The method comprises the following steps: S1, real-time collection of real scene data surrounding the user through the camera and sensors of the mobile device, including scene images, location information, spatial structure, and lighting conditions; S2, semantic segmentation and object recognition of the real scene data using a deep learning model to identify key elements in the scene, including human bodies, clothing stores, fitting rooms, and public areas, and extract scene features; S3, based on the recognition results of the real scene data, obtain user attribute data including gender, age, body type, and real-time behavior data including gaze direction and hand gestures; S4, dynamically select or generate clothing advertisement content from the advertisement database according to the recognition results of the real scene data and the user attribute data, wherein the clothing advertisement content includes 3D clothing models, video advertisements, interactive fitting interfaces, and promotional information; S5, construction of a virtual reality scene on the mobile device, seamlessly superimposing the clothing advertisement content into the real scene through augmented reality or mixed reality technology, wherein the construction of the virtual reality scene includes adjusting the material and shadow of the virtual clothing according to the lighting of the real scene, and using SLAM technology to achieve stable positioning of virtual objects in the real scene; S6, publishing the clothing advertisement content in the virtual reality scene, associating the advertisement with specific objects in the real scene, and displaying it in real time through the mobile device screen; S7, allowing the user to interact with the clothing advertisement content through gestures, voice, or touch screen, including rotating the clothing model, virtual fitting, and obtaining a purchase link; S8, real-time collection of user interaction data and scene data, evaluation of advertisement effectiveness, and dynamic optimization of advertisement content, location, and display timing using machine learning algorithms; S9, providing advertisement publishing management functions through the management interface of the mobile user terminal, including customizing advertisement content, specifying target scenes, setting budget and delivery rules, real-time monitoring of advertisement performance, and generating analysis reports. 2.The method of claim 1, wherein the method further comprises: receiving a request for a virtual reality scene advertisement from a user; and providing the virtual reality scene advertisement to the user based on the request. The use of a deep learning model for semantic segmentation and object recognition of real scene data includes using a convolutional neural network for image analysis to identify clothing-related objects in the scene, and combining sensor data fusion to improve recognition accuracy. 3.The method of claim 1, wherein, The construction of a virtual reality scene includes dynamically adjusting the display angle and size of the virtual advertisement according to the user's location to optimize the viewing experience, and ensuring the visual consistency of the virtual clothing with the real scene through environmental lighting matching. 4.The method of claim 1, wherein, The dynamic selection or generation of clothing advertisement content from the advertisement database includes using a recommendation algorithm to preferentially display clothing styles that match the user's body type and preferences based on user attributes and real-time behavior data. 5.The method of claim 1, wherein, Step S6 includes virtual fitting through gesture recognition, where the mobile device camera captures user gestures and renders the clothing model onto the user's body in real time in the virtual reality scene. 6.The method of claim 1, wherein, The evaluation of advertisement effectiveness includes calculating the click-through rate, dwell time, and number of fittings, and using a reinforcement learning model to adjust the advertisement strategy based on scene complexity data. 7.The method of mobile real scene cognition based virtual reality scene advertisement publishing according to claim 1, characterized in that: The management interface of the mobile user terminal is provided through a graphical user interface, allowing advertisers to upload 3D clothing models and set promotional activities, and remotely update advertisement content based on real-time data.
8. System for advertising in a virtual reality scene based on mobile real scene awareness, based on the method for advertising in a virtual reality scene based on mobile real scene awareness according to any one of claims 1 to 7, characterized in that, The mobile user terminal comprises a real scene cognition module, a virtual reality rendering module, an advertisement interaction module and a publishing management interface, and is used for collecting real scene data, generating a virtual reality scene, displaying clothing advertisement content and providing user interaction and management functions. The server side comprises an advertisement database, a scene analysis engine, an advertisement scheduling engine and a user management module, and is used for storing advertisement content, analyzing real scene data, scheduling advertisement publishing and managing user data. A network communication module is used for realizing data exchange between the mobile user terminal and the server side, and supporting low-delay transmission. The real scene cognition module uses a deep learning model to perform semantic segmentation and object recognition on real scene data, the virtual reality rendering module uses a graphics rendering engine to superimpose clothing advertisement content onto a real scene, and the advertisement scheduling engine selects advertisement content based on real scene cognition results and user attributes. 9.The system for mobile real scene cognition based virtual reality scene advertisement publishing according to claim 8, characterized in that: The mobile user terminal is a smartphone or a tablet computer, which is equipped with a camera, a GPS sensor, a gyroscope and an accelerometer, and the real scene cognition module executes a deep learning model through a built-in processor to identify scene elements in real time. 10.The system for mobile real scene cognition based virtual reality scene advertisement publishing according to claim 8, characterized in that: The advertisement database of the server side stores 3D clothing models and metadata, the scene analysis engine uses a machine learning algorithm to perform advanced scene analysis and returns advertisement recommendation results, and the advertisement scheduling engine is integrated with a rule engine to optimize advertisement publishing according to target scenes and delivery rules set by advertisers.