Method of providing commodity search result information and electronic device

By reconstructing physical spaces in three dimensions and displaying virtual scenes, combined with AR/VR technology, the problem of limited product search methods in existing technologies has been solved, resulting in a richer product search experience and greater interactivity, forming a closed-loop service that integrates online and offline channels.

CN115599991BActive Publication Date: 2026-06-26TAOBAO CHINA SOFTWARE
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TAOBAO CHINA SOFTWARE
Filing Date
2022-09-22
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Current product search methods are limited, relying mainly on keywords or image uploads, lacking diverse interactive features, and failing to provide a more intuitive and richer product search experience.

Method used

Virtual space scenes are generated by shooting real-world locations, and three-dimensional reconstruction of physical spaces is carried out using robotic equipment. The virtual space scenes are then displayed using VR technology, and information cards are displayed through holographic projection or AR methods, which can identify the main body and provide product information in real time.

Benefits of technology

It enables users to have a rich product search experience in a virtual space, provides a new way of "shopping online", enhances user interactivity and the convenience of product search, and forms a closed-loop service between online and offline.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115599991B_ABST
    Figure CN115599991B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method for providing commodity search result information and an electronic device, the method comprising: in response to a request for browsing commodity information in a virtual reality (VR) mode, displaying a pre-generated virtual space scene based on VR technology, the virtual space scene being generated by: taking a real scene of commodity display in an entity space, and performing three-dimensional reconstruction on a space structure of the entity space, splicing and matching the real scene image into the three-dimensionally reconstructed virtual space structure to generate the virtual space scene; performing subject recognition from a real scene image frame corresponding to a current display perspective in the virtual space scene; obtaining a commodity information search result related to the recognized subject; generating an information card to be displayed according to the commodity information search result, and mapping the information card to the virtual space scene for display. Through the embodiments of the present application, a richer commodity search mode can be provided for a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of product search technology, and in particular to methods and electronic devices for providing product search results information. Background Technology

[0002] In traditional product information service systems, product search is based on keywords. Users input product keywords as search criteria, and the system then provides a page of search results matching those criteria. With the development of mobile devices and other technologies, some product information service systems also offer image-based product search functionality. Users can take photos of items they find interesting in real life, or save images they find online if they are interested in the items depicted. These saved photos or images can then be uploaded to the server. Alternatively, users can take photos directly through the photo-taking interface on the client-side page and upload them to the server. Upon receiving the photos or images, the server can identify and compare the subjects in the images to determine and return relevant product search results. Users can then find the products they need from the search results list and even make online purchases.

[0003] In other words, in existing technologies, product searches are conducted either by entering keywords or by uploading / taking photos. How to provide users with more diverse product search methods has become a technical problem that needs to be solved by those skilled in the art. Summary of the Invention

[0004] This application provides a method and electronic device for providing product search results information, which can provide users with richer ways to search for products.

[0005] This application provides the following solution:

[0006] A method for providing product search result information includes:

[0007] In response to a request to browse product information through virtual reality (VR) mode, a pre-generated virtual space scene is displayed based on VR technology. The virtual space scene is generated in the following way: real-scene photography of the physical product display in the physical space, three-dimensional reconstruction of the spatial structure of the physical space, and stitching and matching the real-scene photography images into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0008] Subject identification is performed from the real-scene image frame corresponding to the current display perspective in the virtual space scene;

[0009] Obtain search results for product information related to the identified entity;

[0010] Information cards to be displayed are generated based on the product information search results, and the information cards are mapped to the virtual space scene for display.

[0011] The real-scene images are obtained by robotic equipment in the physical space taking real-scene photos of the physical merchandise display in the physical space.

[0012] The physical space contains multiple different subspaces, and the physical goods include physical goods displayed in the subspaces.

[0013] The real-scene images are obtained by the robotic equipment in the physical space taking real-scene photos of the physical merchandise display in multiple different sub-spaces of the physical space.

[0014] When performing three-dimensional reconstruction of the spatial structure, the spatial distribution structure of multiple subspaces in the physical space is reconstructed in three dimensions.

[0015] This also includes:

[0016] The robotic device provides augmented reality (AR)-based shopping guidance services to users entering the physical space.

[0017] The shopping guide service provided by the robot device includes: real-time image stream acquisition of the physical merchandise display in the physical space, subject recognition from the real-time image stream, obtaining product information search results related to the identified subject, generating information cards to be displayed based on the product information search results, and displaying them.

[0018] Among them, displaying the information card through the robot device includes:

[0019] The information card is mapped to the location of the subject in the corresponding image frame for display.

[0020] Among them, displaying the information card through the robot device includes:

[0021] The information card is displayed using holographic projection.

[0022] When displaying the information card via holographic projection, the method further includes:

[0023] The 3D model of the product corresponding to the information card is displayed using holographic projection.

[0024] The step of performing subject recognition from real-world image frames captured in the virtual space scene includes:

[0025] During the process of dynamically changing the viewpoint and / or performing forward and backward operations in the virtual space scene, subject recognition is performed from multiple real-scene captured image frames, and the state changes of the subject in the multiple real-scene captured image frames are obtained. The state is related to the position of the subject in the real-scene captured image frames and / or the area ratio of the subject image on the screen.

[0026] The step of generating the information card to be displayed based on the product information search results includes:

[0027] Information cards to be displayed are generated based on the product information search results and the status of the subject, wherein different subject statuses correspond to different information card types.

[0028] During the process of the information card moving along with the change in the position of the subject, the latest position of the identified subject is read according to a preset period; wherein the time interval is related to the persistence of vision of the human eye and is greater than the time interval between each frame in the image stream;

[0029] After the latest position is read, an animation is generated for the currently displayed information card to smoothly move from the current position to the latest position, and the position of the information card is kept unchanged at the latest position until a new latest position is read in the next cycle, at which time a new animation is generated.

[0030] The relative positional relationship between the information card and the anchor point of the main body can be varied; the position of the anchor point is determined based on the center point of the main body.

[0031] The method further includes:

[0032] As the information card moves along with the change in the position of the main body, it is determined whether the information card will collide with the edge of the screen in the current relative position relationship. If so, the relative position relationship is switched to another.

[0033] If multiple subjects are identified in the image frame, their respective product information search results are obtained, and information cards of the corresponding type are generated according to their respective states.

[0034] The method further includes:

[0035] During the process of mapping multiple information cards onto the image frame and moving them according to the changes in the position of their respective subjects, it is determined whether a collision will occur between the multiple information cards.

[0036] If so, the display position of the information card corresponding to the subject with the lower priority will be moved according to the different subject status priorities of the information card that collides.

[0037] The information cards corresponding to the different subject states are displayed in different layers, and the layers are arranged in descending order of priority of each subject state.

[0038] This also includes:

[0039] During the process of mapping the information card to the location of the corresponding subject in the image frame for display, if the subject's state changes, the type of information card mapped to the virtual space scene will be switched.

[0040] There is a buffer at the boundary of the judgment parameters for different subject states, so that when the judgment parameters change within the buffer, the current subject state and the type of the corresponding information card remain unchanged.

[0041] The main body's state includes a display state, which includes a first display state and a second display state; in the display state, the information card displays content related to the product search results; the size and content of the information card corresponding to the first display state and the second display state are different.

[0042] The subject's state also includes a discovery state, and the information card in the discovery state is used to display the identification result of the category to which the subject belongs.

[0043] The state of the subject also includes the interactive state;

[0044] The method further includes:

[0045] After the subject enters the interactive state and maintains it for a threshold time, it will be redirected to the product search results list page or the details page of the product included in the currently displayed information card.

[0046] When displaying the virtual space scene using VR or mixed reality (MR) glasses, the method further includes:

[0047] When the VR or MR glasses detect that the user has grabbed the information card and moved it in the target direction, the product corresponding to the information card is added to the user's associated batch checkout product set.

[0048] A robotic device,

[0049] The robotic equipment is used in physical spaces;

[0050] The robot device is used to provide AR-based shopping guide services to users entering the physical space. The shopping guide service includes: real-time image stream acquisition of the physical merchandise display in the physical space, subject recognition from the real-time image stream, obtaining product information search results related to the identified subject, generating and displaying information cards based on the product information search results.

[0051] The robot equipment is also used to take real-scene photos of the physical merchandise display in the physical space, so as to stitch and match the real-scene photos into the three-dimensional reconstructed virtual space structure by performing three-dimensional reconstruction of the spatial structure of the physical space, thereby generating a virtual space scene.

[0052] In the process of VR displaying the virtual space scene through the user's associated terminal device, subject recognition is performed from the real-scene image frame corresponding to the current display perspective, search results for product information related to the identified subject are obtained, information cards to be displayed are generated based on the search results for product information, and the information cards are mapped to the virtual space scene for display.

[0053] An information display method, comprising:

[0054] The virtual space scene is generated by displaying a pre-generated virtual space scene based on VR technology. The virtual space scene is generated by taking real-scene photos of the entities in the physical space and reconstructing the spatial structure of the physical space in three dimensions. The real-scene photos are then stitched and matched into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0055] Subject identification is performed from the real-scene image frame corresponding to the current display perspective in the virtual space scene;

[0056] Based on the identified subject image, obtain search results for information related to the subject;

[0057] Information cards to be displayed are generated based on the search results, and the information cards are mapped to the virtual space scene for display.

[0058] An apparatus for providing product search results information includes:

[0059] The VR display unit is used to respond to requests to browse product information through virtual reality (VR) mode and display a pre-generated virtual space scene based on VR technology. The virtual space scene is generated in the following way: real-scene shooting of the physical product display in the physical space, three-dimensional reconstruction of the spatial structure of the physical space, and stitching and matching the real-scene shooting images into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0060] The subject recognition unit is used to perform subject recognition from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene;

[0061] The product search result acquisition unit is used to acquire product information search results related to the identified subject;

[0062] The information card display unit is used to generate information cards to be displayed based on the product information search results, and to map the information cards to the virtual space scene for display.

[0063] An information display device, comprising:

[0064] The VR display unit is used to display a pre-generated virtual space scene based on VR technology. The virtual space scene is generated in the following way: real-scene shooting of entities in the physical space, three-dimensional reconstruction of the spatial structure of the physical space, and stitching and matching the real-scene shooting images into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0065] The subject recognition unit is used to perform subject recognition from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene;

[0066] The search unit is used to obtain search results for information related to the identified subject image.

[0067] The information card display unit is used to generate information cards to be displayed based on the information search results, and to map the current display perspective onto the virtual space scene for display.

[0068] A computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any of the preceding methods.

[0069] An electronic device, comprising:

[0070] One or more processors; and

[0071] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the preceding descriptions.

[0072] According to the specific embodiments provided in this application, the following technical effects are disclosed:

[0073] Through the embodiments of this application, the actual product display in a physical space can be photographed in advance, and the spatial structure of the physical space can be reconstructed in three dimensions. The photographed images are then stitched and matched into the reconstructed virtual space structure to generate a virtual space scene. When a user requests to browse product information in VR mode, the pre-generated virtual space scene can be displayed using VR technology, and subject recognition can be performed from the image frames within the virtual space scene. Then, product information search results related to the identified subject can be obtained, and information cards to be displayed can be generated based on these search results and mapped to the corresponding position of the subject in the image frame. In this way, users can experience "online shopping," and the photographed images within the virtual space scene can serve as the starting point for "image-based product search." Through subject recognition, product search results are automatically obtained and mapped back to the virtual space scene in the form of information cards for display. This makes the information in the virtual space scene created through photographed images richer and provides users with a completely new product search and interaction experience.

[0074] Furthermore, to facilitate the creation, updating, and maintenance of virtual space scenes, this application embodiment also provides a robotic device. This device can capture real-world images of merchandise displays in physical spaces to generate virtual space scenes, which can be updated daily. Additionally, the robotic device can provide shopping guidance services to users entering physical spaces. This includes recognizing objects while collecting images of physical merchandise displayed in the mall, generating information cards from search results, and then mapping these cards onto the image stream using AR technology. In this way, a closed loop of "online shopping" and "offline shopping" can be formed using the robotic device, providing users with image-based product search services both online and offline.

[0075] Furthermore, whether "shopping online" through a user terminal device or "shopping offline" through a robot device, the image frames are dynamically changing, causing the position and state of the main image within the image frame to change accordingly. Therefore, in this embodiment, different types of information cards can be provided for different main states, and the information cards can follow the movement of the main body.

[0076] Regarding the issue of discontinuous motion trajectories encountered by information cards during the movement of the main body, a preferred embodiment of this application can solve this problem through segmented fitting. Regarding collisions between information cards and screen edges, two relative positional relationships between the information card and the center point of the main body can be pre-defined; when such collisions occur, the relative positional relationship can be switched to resolve the issue. Regarding collisions between information cards, the priority of different types of information cards can be determined based on the priority relationship between different main body states, thus identifying lower-priority information cards as those that need to be moved. Alternatively, multiple layers can be created, with information cards corresponding to different main body states displayed on different layers.

[0077] Furthermore, in XR glasses scenarios, the interactive method of adding products to the shopping cart by having users "reach out and grab" information cards can be combined to further enhance the fun of the interaction and user participation.

[0078] Of course, any product implementing this application does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0079] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 This is a schematic diagram of the system architecture provided in the embodiments of this application;

[0081] Figure 2 This is a flowchart of the first method provided in the embodiments of this application;

[0082] Figure 3 This is a schematic diagram of the display interface provided in an embodiment of this application;

[0083] Figure 4 This is a schematic diagram of the timeline provided in an embodiment of this application;

[0084] Figure 5 This is a flowchart of the second method provided in the embodiments of this application;

[0085] Figure 6 This is a schematic diagram of the first device provided in the embodiments of this application;

[0086] Figure 7This is a schematic diagram of the second device provided in the embodiments of this application;

[0087] Figure 8 This is a schematic diagram of the electronic device provided in the embodiments of this application. Detailed Implementation

[0088] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0089] In this embodiment, a new product search method can be provided to users. This method also allows for image-based product searches, but without requiring users to upload or take photos / videos. Instead, the system can pre-create a virtual space scene based on the product display in a physical location (e.g., a shopping mall). This virtual space scene can be created by taking real-world photos of the product display in the physical location, reconstructing the spatial structure of the physical location in 3D, and then stitching the real-world images into the reconstructed virtual space structure. In other words, in the virtual space scene created in this embodiment, it is not necessary to create virtual 3D models for the products within the scene; instead, images taken in real-world locations are used to represent them. This makes the virtual space scene more realistic and reduces maintenance costs.

[0090] Furthermore, given the aforementioned virtual space scenario, this application embodiment can provide a function for searching product information based on real-world images captured within this virtual space scenario. Specifically, when a user browses the virtual space scenario using VR (Virtual Reality), the image presented to the user is an image from a certain perspective within the virtual space scenario, which is composed of real-world images. Of course, the user can switch perspectives by swiping the screen to view images from other perspectives within the virtual space scenario. During this process, subject recognition can be performed on the image frames (where the subject in image recognition typically refers to the photographed object located in the foreground of the image; in this application embodiment, the specific subject is usually an image of a physical product), and relevant product information search results can be obtained based on the identified subject. Then, information cards to be displayed can be generated based on the product information search results and mapped onto the image frames for display. For example, the specific information card can display the price, discount information, etc., of the searched product.

[0091] In other words, in this embodiment, users can experience "online shopping" through a virtual space. They can not only view the product displays in physical locations like shopping malls, but also access additional information about the products through displayed information cards. Furthermore, these real-world images, presented in a virtual space, also serve as an entry point for image-based product searches. Users can "shop online" while simultaneously obtaining product search results through "image search," thus providing a completely new way to initiate and experience "image search." For merchants, this can transform customer traffic from "one city" to "the entire internet," significantly boosting sales.

[0092] Specifically, the physical space can be a shopping mall or similar venue. When creating the corresponding virtual space scene, since it is necessary to take real-scene photos of the display of physical goods in the physical space, in a preferred embodiment of this application, a robot device can also be provided. During non-business hours of shopping malls or similar venues, the robot device can "walk" in the mall and enter each store to take real-scene photos of the goods display. Then, the server can generate a specific virtual space scene based on the real-scene photos uploaded by the robot.

[0093] In addition, the robot can provide shopping guidance services to consumers during the mall's business hours. During the shopping guidance service, it can also collect real-time image streams of physical products displayed in the mall, and can also perform subject recognition based on the real-time image streams, and search for products based on the recognized subjects. Afterwards, it can also generate information cards, which can be displayed through AR or holographic projection.

[0094] In this way, a closed loop of online and offline "shopping" can be achieved through robotic devices. That is, while users are shopping in a physical store, the robotic devices can search for additional information about the products they actually see and generate information cards, which are then displayed to the user through AR or holographic projection. Furthermore, virtual space scenes can be created using the real-world images captured by the robotic devices, providing users with a "shopping online" experience through VR. Moreover, during this "shopping online" process, users' devices can also search for additional information about the products they see in the virtual space scene, generate information cards, and display them within the VR visuals.

[0095] Furthermore, whether information cards are displayed via AR using robotic devices or via VR virtual space scenes displayed on user terminal devices, the image stream captured by the robotic devices dynamically changes with the movement of the robotic devices. The virtual space scene displayed on the terminal device can also change its perspective, move forward or backward, by swiping the screen or rotating the terminal device. Therefore, the images of the virtual space scene displayed on the terminal device are also dynamically changing. Correspondingly, the position and state (including the size of the subject image) of the identified subject in the image frame also change in real time. Therefore, in this embodiment, the position and state of the subject can be recorded in real time, and various types of information cards can be predefined. Different types of information cards can be displayed under different subject states. Specifically, the subject state can be determined based on the subject's position in the image, the area ratio between the subject image and the screen, etc., with the specific area ratio varying according to the subject's distance from the camera. For example, the subject state can be divided into discovery state, display state, and interactive state. The display state can be further subdivided into first display state, second display state, and so on.

[0096] Furthermore, in the scenario where information cards are mapped onto image frames, and multiple different types of information cards correspond to different subject states, processing of information card display is required on the client side (robot device side or user terminal device side), which may present some challenges. For example, if the information card needs to move with the subject, how can the movement trajectory of the information card be made smoother and more fluid, provided the client-side performance can support it? Alternatively, since the information card needs to occupy a certain area, it may collide with the screen edge during its movement. When multiple subjects are identified in the image frame, multiple information cards need to be mapped onto the image frame, and collisions may also occur between these cards. When such collisions occur, the content of the information card may be obscured by the screen edge or other information cards; how to solve this problem also needs to be considered. Additionally, since different types of information cards need to be switched when the subject state changes, and changes in subject state are usually caused by changes in the distance between the subject and the lens, or by changes in the viewing angle, if the distance or viewing angle is exactly at the boundary between two subject states, jitter may occur. In other words, information cards may repeatedly and rapidly switch between two types, and so on. This application provides corresponding solutions to each of the above problems in real time.

[0097] In addition to controlling the movement and switching of the information card as described above, in practical implementation, some "action points" can be pre-emptively implemented based on this information card. For example, a specific "action point" could include adding a product to the "shopping cart." That is, users can directly add products to their cart through this information card without needing to navigate to the product search results page or product details page. Specifically, if implemented on mobile devices such as smartphones, the information card can provide specific options for adding the displayed products to the "shopping cart," allowing users to complete the addition by clicking on these options.

[0098] Alternatively, when "shopping online" using VR or MR (Mixed Reality) glasses, users can add items displayed on information cards to their cart through richer interactive methods. For example, users can use the controllers of their MR glasses or make certain gestures to grab information cards and move them in a specific direction (e.g., to their chest), thereby adding the displayed items to their cart.

[0099] From a system architecture perspective, see Figure 1This application embodiment may involve the client and server sides of a commodity information service system. The client in this application embodiment may include an App (application) installed on a mobile phone or other terminal device, or an App installed on VR / MR glasses, etc. Furthermore, since it also involves real-scene photography during the virtual space scene creation process, in an optional implementation, a robot device may be provided to take real-scene photos of the commodity display in a physical space (which may include multiple sub-spaces, such as multiple specific physical stores). The server can create a virtual space scene based on the real-scene photos, and the client can display the virtual space scene via VR, perform subject recognition from the real-scene photo frames, obtain commodity search results from the server, and generate information cards to be mapped back to the virtual space scene for display. Additionally, the robot device can provide shopping guidance services to users entering the physical space, and provide more additional information about the products in the mall during the shopping guidance process using AR. Specifically, algorithms related to subject recognition and search request initiation can be deployed on the client side (including clients on the robot device side and user terminal devices). Furthermore, the display control of information cards during movement or switching, including image stream processing, card rendering, subject following, card collision avoidance, and card switching, can also be handled on the client side. The server primarily creates virtual space scenes and provides specific product search results, which the client can then use to generate specific information cards.

[0100] The specific implementation schemes provided in the embodiments of this application will be described in detail below.

[0101] Example 1

[0102] First, from the perspective of the user terminal device side client in this embodiment, a method for providing product search result information is provided, see [link to previous document]. Figure 2 The method may include:

[0103] S201: In response to a request to browse product information through virtual reality (VR) mode, a pre-generated virtual space scene is displayed based on VR technology. The virtual space scene is generated by: taking real-scene photos of the physical product display in the physical space, reconstructing the spatial structure of the physical space in three dimensions, and stitching the real-scene photos into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0104] The physical space can be a shopping mall, or even an independent shop, etc. This embodiment of the application generates the virtual space scene by taking real-world photos of the merchandise display in such a physical space, reconstructing the spatial structure of the physical space in 3D, and then stitching the real-world images into the reconstructed virtual space structure. In other words, this embodiment generates the virtual space scene by stitching together real-world images, without needing to perform 3D modeling of the physical merchandise in the physical space. Therefore, the cost is lower, and the result is more realistic.

[0105] After creating a virtual space scene, it can be published online. For example, if a physical location corresponds to an online store, an entry point for browsing via VR can be provided on the store's homepage or other pages. Users can then use this entry point to request to browse product information via VR. Alternatively, a "shopping mall online" function module can be provided in the client application, with an entry point for this module on the client's homepage or other pages. On the specific landing page, access points for multiple virtual space scenes can be centrally displayed. Users can then select a specific virtual space scene while accessing this landing page and request to browse product information within it via VR, and so on.

[0106] Regarding the creation of virtual space scenes, one approach is to manually photograph the product display in a physical space and upload the images to a server, which then generates the corresponding virtual space scene. Alternatively, considering that the product display in a physical space may frequently change—for example, a shop might change the items displayed in its window daily—it may be necessary to frequently re-photograph the physical scene and update the virtual space scene. Therefore, in another approach, this application embodiment also provides a robotic device that can photograph the physical product display in a physical space during non-business hours to obtain specific real-scene images.

[0107] In the case of physical spaces such as shopping malls, which typically contain multiple sub-spaces (e.g., physical stores), the physical goods displayed in the physical space include those within specific physical stores. Therefore, real-scene filming can be conducted during non-business hours of the physical space, using robotic equipment to film the product displays in multiple different physical stores. In other words, robotic equipment, or robots with machine vision, can be controlled to "walk" through the mall and enter each physical store to film the product displays, uploading the footage to a server. Individual physical stores can be pre-positioned with display booths to support this real-scene filming. The product display method in these booths can also follow specified standards; for example, different products can be spaced a certain distance (to reduce collisions between different information cards during subsequent VR demonstrations), and for clothing, the front of the product needs to be captured (for subject identification and accurate search results), etc.

[0108] In the case of multiple physical stores within the same physical space, when performing three-dimensional reconstruction of the spatial structure, the spatial distribution structure of the multiple physical stores in the physical space can be reconstructed in three dimensions to better restore the real scene of the shopping mall.

[0109] The specific details of how to conduct real-scene shooting and how to construct virtual space scenes based on real-scene images are not the focus of this application's embodiments and will not be elaborated here.

[0110] S202: Perform subject identification from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene.

[0111] In step S201 above, a virtual space scene is displayed using VR technology. During the display of this virtual space scene, subject recognition can be performed on the image frames. The so-called real-scene captured image frame refers to the fact that the virtual space scene is usually a 3D scene, while the user terminal screen is usually rectangular. Therefore, only a portion of the image in the virtual space can be viewed at any given time. Users can switch perspectives by swiping the screen or rotating the terminal device, or by clicking forward or backward buttons in the scene to move forward or backward within the virtual space. The visual effect is similar to that of dynamically capturing an image stream through a camera. Users can browse multiple frames, each of which can be called an image frame. Furthermore, since the images in the virtual space scene in this embodiment are stitched together from real-scene captured images, each image frame can be called a real-scene captured image frame.

[0112] Subject recognition primarily involves identifying foreground objects within an image frame. For instance, if an image frame was captured by photographing goods displayed in a shop, those goods can be identified as the subject. While the background of the shop is also captured, it won't be recognized as the subject. Of course, multiple subjects may be identified within the same image frame; for example, multiple goods entering the shooting range simultaneously can be identified as multiple subjects. Specific subject recognition results can include the subject's ID, coordinates, dimensions, category, and whether a search request can be sent to the server. The client can record these recognition results for subsequent use in generating and displaying information cards, as well as in motion / switching control.

[0113] The subject number is primarily used to label the specifically identified subjects. This label is provided by the algorithm, and the same subject has the same number in each frame's recognition result. Thus, when multiple subjects are identified in an image frame, they can be distinguished by their subject numbers, enabling subsequent information card tracking and other functions.

[0114] The subject's coordinates primarily refer to the coordinates of its center point on the screen, while its length and width refer to the maximum width and maximum length of the subject image on the screen. During dynamic changes in viewpoint, the coordinates, length, and width of the same subject are constantly changing, which can be identified using the aforementioned client-side algorithm. Correspondingly, this information can be recorded by the client.

[0115] The main category refers to the product category to which the subject image likely belongs; that is, identifying the specific item the subject is, such as a cup, headphones, etc. This information can be used to generate a "catch-all" information card. In other words, when a specific product search result cannot be obtained, information such as the category name can be displayed in the information card based on the product category to which the subject belongs.

[0116] Regarding whether to initiate a search request to the server, specific product search results can be obtained from the server. However, because the size and clarity of the main image constantly change during dynamic perspective changes or forward / backward operations in a virtual space scene, when the main image is relatively small, the server may find it difficult to provide accurate product search results based on such an image. Therefore, it may not be necessary to submit a search request to the server. Later, if the area of ​​the main image is large enough and clear enough, it can be determined whether to initiate a product search request to the server. Furthermore, as the main image becomes clearer, the server may be able to provide more accurate search results; therefore, a new product search request can be initiated, and so on. All of these decisions can be made by the aforementioned client-side algorithm during the subject recognition process.

[0117] After the algorithm identifies the above information, it can be recorded on the client side, and the location information of the subject can be determined based on information such as the coordinates of the subject's center point.

[0118] In practical applications, when VR technology is used to display virtual space scenes, users can change their viewing angle or perform actions such as forward and backward by swiping the screen or rotating the terminal device. The image displayed on the terminal device screen will also dynamically change with the change of viewing angle (similar to the presentation effect of real-time image stream acquisition through a camera). Therefore, the process of subject recognition from image frames is also dynamic and continuous. When recognition is performed at a specific moment, it is from the image frame currently displayed to the user.

[0119] Furthermore, in practical implementation, during the process of dynamically changing the viewpoint or performing forward and backward operations in the virtual space scene, subject recognition can be performed from multiple image frames, and the changes in the subject's position and state across these frames can be obtained. This allows for the generation of information cards to be displayed based on product information search results and the subject's state, with different subject states corresponding to different information card types.

[0120] Among them, regarding the subject state, that is, the position of the subject in the image frame and / or the proportion of the subject image in the area of the screen, etc. For example, in one way, the subject state can be determined simply based on the area ratio of the subject image in the screen. Specifically, for example, when the subject state is divided into discovery state, first display state, second display state, and interaction state, when the area ratio of the subject image in the screen (assumed to be represented by the letter M) is very small, for example, a < M < b. At this time, it may only be possible to identify what category of item the subject is, but it is impossible to determine which actual sellable products in the system it is related to. Therefore, it can be determined that the subject is in the "discovery state". After that, as the area of the subject image in the screen gradually increases, for example, b ≤ M < c, the image also gradually becomes clear. When it reaches the level where product search can be performed based on the subject image, a search can be initiated to the server. Correspondingly, the subject enters the first display state. After that, when the area of the subject image further increases, for example, c ≤ M < d, and the image is even clearer, a search request can also be initiated to the server again (or, it may not be necessary to request again, depending on actual needs). At this time, the subject can enter the second display state. After that, if the area of the subject image further increases, for example, d ≤ M < e, and lasts for a certain period of time (for example, 3S), then the subject enters the interaction state, and so on.

[0121] Or, in another case, when determining the subject state, in addition to considering the factor of the area ratio of the subject image, the coordinate factor of the subject image can also be considered. For example, regarding the multiple predefined subject states, in addition to defining the area ratio of the subject image corresponding to each subject state, the position information corresponding to each subject state can also be defined. For example, for the discovery state, first display state, second display state, and interaction state, "discovery frame", "display frame", "interaction frame", etc. can be predefined. In this way, only when the area ratio of the subject image is a < M < b, and the center point of the subject is inside the "discovery frame" and outside the "display frame", can it be determined that the subject is in the discovery state. Similarly, only when the area ratio of the subject image is b ≤ M < c or c ≤ M < d, and the center point of the subject is inside the "display frame" and outside the "interaction frame", can it be determined that the subject is in the display state. If the area ratio of the subject image is d ≤ M < e, and the center point of the subject is inside the "interaction frame", it can be determined that the subject is in the interaction state, and so on.

[0122] S203: Obtain the search result of product information related to the identified subject.

[0123] After identifying the subject from the image frame, search results for related product information can be obtained. Specifically, as mentioned earlier, relevant search results can be requested from the server, or, if the terminal device has a local cache, they can be read from the cache, and so on. Specifically, when searching through the server, the subject image can be extracted from the image frame, and a search request can be submitted to the server, which will then retrieve the search results that meet the specified criteria.

[0124] Furthermore, since the client-side algorithm can also identify the subject's ID, category, etc., the identification results regarding the subject's ID, category, etc., can also be submitted to the server when submitting a search request. The server can then use this information to search for product results related to the subject's image. The returned product search results can also include specific subject ID and other information.

[0125] It should be noted that, in this embodiment, relevant merchants can pre-publish information about the goods displayed in their physical locations to the server for storage. This allows the server to directly search for goods corresponding to the physical goods displayed in the physical location when searching based on the main image identified from the real-scene photograph. Of course, in practice, the server's product database may also contain similar or identical products; therefore, this product information can also be provided in the search results for user comparison.

[0126] S204: Generate an information card to be displayed based on the product information search results, and map the information card onto the image frame for display.

[0127] After obtaining product information search results, information cards to be displayed can be generated based on these results. These information cards primarily display text information, such as product names, prices, and user reviews obtained from the search results. This information can then be mapped to a virtual space scene for display. Preferably, it can be mapped to the location of the corresponding subject within the virtual space scene. This allows users to not only browse the products included in the virtual space scene via VR but also obtain additional information about specific products from these information cards. It should be noted that while the virtual space scene in this embodiment is generated by stitching together real-world images, which offers advantages such as lower maintenance costs and higher realism compared to 3D model reconstruction of products, it suffers from insufficient information richness. In this embodiment, however, by performing subject recognition based on real-world images within the virtual space scene and then mapping product search results to the virtual space scene using information cards, information richness is improved. This allows users to obtain a more realistic spatial scene while acquiring richer information.

[0128] As mentioned earlier, during the display of a virtual space scene, users can change their viewing angle, causing the image content displayed on the screen to change dynamically. Consequently, the identified subject's position and state will also change accordingly. Therefore, in a preferred embodiment of this application, different types of information cards can be provided based on different subject states.

[0129] Regarding the subject's state, as mentioned earlier, it can include discovery state, display state, and interactive state. The types of information cards can be categorized primarily by card size and content richness. For each different subject state, the corresponding information card type can be pre-defined. For example, when the subject first enters the display state, the content in the information card can come from product search results; however, the information card can be relatively small, and the amount of information displayed may be limited. Later, as the subject image increases in size, it can switch to a larger information card, displaying more information.

[0130] In short, information cards can come in various types. As the area occupied by the main image increases, the information card can vary in size, and the information displayed can become more comprehensive. For example, ... Figure 3As shown in (B) (for ease of illustration, only one main image is shown in the figure), assuming the currently identified main object is a pair of headphones, when the main image occupies 40% (or other values), a smaller information card can be displayed. This smaller information card can include only the product name, reference price, representative user reviews, etc. Figure 3 As shown in (C), when the main image is enlarged, for example, when the image occupies 60% (or other values), a larger information card can be displayed. This card can include not only the information shown in the smaller card but also richer information such as product performance details. Furthermore, when displaying information about a specific product in the search results within a specific information card, in addition to the product description information mentioned above, marketing information can also be included. For example, if a merchant has configured a coupon for a product, this information card can be presented to the user, allowing them to know about the specific product's discount information before entering the search results page, thus improving click-through rates, conversion rates, and other metrics. Alternatively, since the subject in the actual real-life image is a physical product, and may be a product the user has already purchased, merchants can also configure user benefits such as "buy one get one free" for certain special products (e.g., beverages). In this case, information about these user benefits can be provided through an information card, which users can click to obtain, thereby increasing interactivity, and so on.

[0131] Of course, as mentioned earlier, in practical implementation, the identified subject can also include the discovery state. That is, at this point, the area of ​​the subject image is very small, insufficient to accurately provide a precise product search result, but the category of the object can be roughly identified. Therefore, a corresponding information card can also be generated based on this category information. For example, ... Figure 3 As shown in (A), recognition begins when the subject image occupies 20% of the total image area. However, at this point, the displayed information card may only show the object category information, and so on. Of course, in practical applications, when the subject is in the detection state, the initially identified object category may be incorrect. Subsequently, as the subject image gradually increases in size, the category recognition result can be corrected, and the content displayed in the information card can be updated.

[0132] In addition, the identified subject can also include the interactive state. In this case, the subject image occupies a large area, the image is clear enough, and an information card for the interactive state can be displayed. Alternatively, the information display format can be changed. For example, instead of displaying information cards, a 3D model of the product can be mapped onto a real-time information stream. There are various specific interaction methods. For instance, one approach is to redirect to a search results list page after entering and maintaining an interactive state for a certain period (e.g., 3 seconds), or even redirect to the details page of the currently displayed product, and so on.

[0133] It's worth noting that while displaying additional information about specific products through information cards, these cards can also provide "action points," such as adding the product to a bulk checkout cart. Specifically, on mobile devices like smartphones, the information card can offer an option to add the corresponding product to the cart. This allows users interested in a product to directly add it to their cart by clicking this option, without needing to navigate to the search results page and then the product details page, thus shortening the user's workflow.

[0134] Alternatively, the solution provided in this application can be used not only in mobile terminals such as smartphones, but also in VR or MR glasses devices. That is, the functions provided in this application embodiment can be implemented in related applications on the VR or MR glasses device. In this way, when a user is "browsing an online shopping mall" while wearing VR or MR glasses, one advantage of VR or MR glasses over mobile terminals such as smartphones is that VR or MR glasses do not require handheld use, freeing the user's hands. Simultaneously, the hands can perform some control operations on the screen content through the controller associated with the VR or MR glasses, or by making some gestures, etc. Therefore, in this VR or MR glasses scenario, the operation of adding items to the "shopping cart" does not need to provide a corresponding add-to-cart button on the information card. Instead, it can be achieved by detecting the user's grasping action on the information card and moving it in a target direction (e.g., towards the user's chest, etc.). In other words, when viewing products and information cards in a virtual space scene through VR or MR glasses, if a user is interested in a product presented on an information card, they can use a controller or gestures to grab the information card and move it towards their chest, thus triggering the "add to cart" action for that product. This interactive method of "reaching out to grab" the information card further enhances the fun of the interaction and user engagement.

[0135] Additionally, it should be noted that when robotic devices are deployed in physical spaces, besides capturing real-time images of product displays, they can also provide shopping guidance services to customers entering the space. Specifically, during the physical space's operating hours, the robotic devices can provide AR-based shopping guidance services. This service may include: real-time image streaming of product displays, subject recognition from the image stream, obtaining product search results related to the identified subjects, generating and displaying information cards based on the search results.

[0136] In other words, after entering the aforementioned physical space, users can browse the mall guided by these robotic devices. These devices not only help users quickly locate the specific shops they wish to visit, but also use AR technology to search for more information about the products displayed in the mall online, aiding in their shopping decisions. For example, when a user enters a shop, they can activate the robot's AR function. At this point, the robot can turn on its camera to capture a real-time image stream and perform subject recognition during the image stream acquisition process.

[0137] The subject recognition process is similar to the aforementioned process of subject recognition from real-world images captured in VR scenes. It can also dynamically determine the subject's position and state, and generate different types of information cards for display. It's important to note that the robot device can display these information cards in various ways. For example, one method is to use AR (Augmented Reality) to map the information card to the location of the subject in the corresponding image frame. That is, the user can view an AR screen on the robot device's display, which includes the real-time captured image and the aforementioned information card.

[0138] Alternatively, the robot device's client can display the information cards via holographic projection. That is, users can view the specific information cards from outside the robot device's display screen. Furthermore, in this method, 3D models of the products corresponding to the information cards can also be displayed via holographic projection. In other words, users can experience the product "standing up" in front of them, and the 3D model can even rotate automatically, allowing users to view the product from multiple perspectives. Simultaneously, users can obtain additional information about the specific product through the information cards displayed nearby, including the product name, price, user reviews, and so on.

[0139] In the process of displaying information cards using robotic devices via AR or holographic projection, there is also an option to add the corresponding products to a bulk checkout cart. Users can then click this option to add items to their cart, and so on.

[0140] In addition to providing shopping guidance services, the aforementioned robotic devices can also offer users information such as reference opinions. For example, when a user is trying on or wearing a product, they can ask the robotic device "Does it look good?" The robotic device can use machine vision and related algorithms to evaluate the user's experience of trying on or wearing the product. It can also provide comparative opinions on similar products that the user has tried on or worn during the current shopping process, and so on.

[0141] The above describes implementation schemes for image subject recognition in both online and offline shopping malls, using virtual space scenes or robotic devices, and providing information cards based on product search results derived from the subject image. In one approach, after generating a specific information card, it can be mapped to the location of the subject in the image frame (the image frame currently displayed on the screen in the virtual space scene, or the current image frame in the image stream captured by the robotic device). However, in practice, since the subject's position is constantly changing, the information card needs to move with the subject. During this movement, the information card may collide with screen edges or other information cards. Furthermore, when the subject's state changes, it involves switching the information card type, and so on. This application proposes solutions to address these issues.

[0142] In the process of moving the information card along with the subject, the algorithm can perform subject recognition for each frame in the image stream. The subject's position may change in each frame, and the identified subject positions are multiple discrete points. Therefore, if the information card's position is directly updated based on the identified subject positions in each frame, the information card's movement trajectory may not be smooth enough. Furthermore, this poses a significant challenge to client performance.

[0143] Therefore, a segmented fitting implementation scheme is adopted in this embodiment. Specifically, the client can save the algorithm recognition results. During the process of the information card moving along with the change in the subject's position, the latest position of the recognized subject can be read at preset time intervals. This time interval can be related to the persistence of vision in the human eye. For example, in a specific implementation, this time interval can be 0.35 seconds, that is, the latest position of the subject is read every 0.35 seconds. After reading the latest position, an animation can be generated for the currently displayed information card to smoothly move from its current position to the latest position. Then, the position of the information card can be kept at this latest position until a new latest position is read in the next cycle, at which point a new animation is generated.

[0144] For example, such as Figure 4The timeline shown illustrates an information card that begins displaying at time A. Initially, based on the main subject's position at time A, the information card is displayed at that position without animation. Then, an animation update is performed every n1 seconds. For example, if the time interval between time A and time B is n1 seconds, although the algorithm updates the main subject's position multiple times, and the information recorded on the client side is also continuously updated, the animation remains unchanged; the information card remains at the position it was at time A. When time B arrives, the most recently updated main subject position is used to initiate the information card's movement. Specifically, this movement could involve a translational animation lasting n1 seconds. That is, the information card smoothly moves from its position at time A to its position at time B along a straight line. Afterward, the information card remains at its position at time B until time C arrives n1 seconds later, at which point it moves again using a translational animation lasting n1 seconds, and so on. In other words, in this method, the movement trajectory of the information card is divided into multiple segments of linear motion. These segments are connected to fit the main movement trajectory, making the movement of the information card smoother and reducing stuttering. Furthermore, the multi-segment translational animation is also more performance-friendly for the client.

[0145] Regarding the issue of information cards colliding with the screen edge, since information cards are usually rectangular, they may collide with the screen edge while following the movement of the main image, resulting in part of the information card's content being obscured and unable to be displayed.

[0146] To address this issue, in this embodiment, the relative positional relationship between the information card and the main anchor point (which can be determined based on the position of the main center point; for example, the main center point can be used directly as the anchor point, or the anchor point and the main center point can be on the same horizontal or vertical line, etc.) can be defined in advance as various types. For example, two different relative positional relationship types can be defined: a first relative positional relationship and a second relative positional relationship. Specifically, they can be lower right and upper left, respectively. That is, when displaying the information card, the information card can only appear in the lower right or upper left position of the anchor point. Thus, as the information card moves with the change in the position of the main body, it can be determined whether the information card will collide with the edge of the screen in the current relative positional relationship state. If so, it can be switched to another relative positional relationship. For example, the information card was originally in the lower right position of the main anchor point, but as the main body moves, the information card may collide with the right edge of the screen, so that part of the information in the information card may be obscured. At this time, the information card can be switched to the upper left position. In determining whether a collision will occur with the screen edge, assuming the current information card is displayed to the lower right of the main anchor point, the main determination is whether the information card will collide with the right edge of the screen. Specifically, the X-axis coordinate of the main center point can be added to the length of the card, and the result can be determined whether it exceeds the right edge of the screen, etc.

[0147] Of course, in actual implementation, if it is found that the information card will exceed the right edge of the screen in the current relative position relationship, it may also exceed the left edge of the screen after switching to other relative positions of the main anchor point. Therefore, before switching, it is also possible to determine whether the information card will exceed the left edge of the screen after switching to other relative positions of the main anchor point. If so, there is no need to switch; otherwise, switch.

[0148] In the aforementioned implementation method of segmented smooth fitting of the motion trajectory of the information card, the latest position of the subject can be read every n1s, and after moving the information card, it can be determined whether the information card will exceed the edge of the screen.

[0149] In addition, such as Figure 3 As shown at point 31 in (B) of the diagram, specific anchor elements can also be displayed on the main image during the specific display of the information card. When switching relative positions, the anchor elements can be hidden first, and then made visible again after the relative position switch is complete. This makes the switching process more aesthetically pleasing.

[0150] The above describes solutions for collisions between information cards and screen edges. In practical applications, multiple subjects may be identified in the same image frame. In this case, if the corresponding product information search results are obtained separately and information cards of the corresponding type are generated according to their respective states, collisions may occur between the multiple information cards during the process of mapping them to the image frame in real time and moving with the position changes of their respective subjects. That is, one information card may obscure the content of another information card.

[0151] To address this situation, this embodiment can further determine whether multiple information cards will collide. If so, the obstruction can be eliminated by moving one or more information cards. Regarding which specific information card to move, this embodiment can adjust the display position of the information card corresponding to the lower-priority subject based on the different priorities of the colliding information cards. In other words, since this embodiment defines multiple different subject states, and subject states are affected by factors such as the size of the subject image, multiple subjects in the same image stream may be in different states. Furthermore, the larger the area of ​​the subject image, the more likely it is to be the object the user most wants to search for; therefore, the corresponding information card can be displayed with a higher priority. This allows for sorting of different subject states according to priority, and when a collision is detected between two information cards, the information card corresponding to the lower-priority subject can be moved.

[0152] It's important to note that if a card is moved, it might collide with the screen edge, causing the edge to obscure the card and preventing further movement. Therefore, to mitigate the impact of collisions between different cards (primarily obscuring content), multiple layers can be created. Each layer displays cards corresponding to a specific theme state, and these layers can be prioritized according to the theme state. For example, the first layer could hold cards for interactive states, the second for display states (large cards), the third for discovery states (small cards), and so on. This way, when a theme is in a particular state, the corresponding card is displayed in the layer for that state. Layer switching can also be performed simultaneously when displaying different types of cards for the same theme. For example, if a subject is previously in the first display state and the information card type is a "small card" for the display state, then the information card can be displayed in the aforementioned third layer; at some point, the subject's state changes to the second display state and the information card type switches to a "large card" for the display state. At this time, the "large card" can be mapped to the second layer in the image stream, and so on.

[0153] In this way, information cards corresponding to different subject states are displayed on different layers, with higher-priority subject states appearing on the uppermost layer. This ensures that if information cards corresponding to two different subject states collide, the higher-priority card remains on the upper layer. Furthermore, even if collisions occur and lower-priority cards cannot be moved, it prevents higher-priority cards from being obscured by lower-priority cards.

[0154] Furthermore, since this application embodiment defines multiple different subject states, different subject states can correspond to the generation of different types of information cards. Moreover, the subject state of the same subject often changes within an image frame. Therefore, the specific display of information cards also involves the switching of information card types. That is, during the process of mapping the information card to the location of the corresponding subject in the real-time image stream for display, if the subject state changes, the type of information card mapped to the real-time image stream will be switched. For example, if a smaller information card is displayed for an entity at a certain moment, and later, the entity's image area increases and the image becomes clearer, then a larger information card can be switched. For example, as... Figure 3 As shown in (B) and (C) in the diagram.

[0155] Among them, since the switching of the information card type is triggered by the change event of the main body state, and the change of the main body state is related to the change of the area ratio of the main body image. For example, as described above, the area ratio is divided into four intervals, and each interval corresponds to a main body state. If a certain determined value is used as the boundary of various main body states, at the boundary of the two main body types, the situation of "jitter" of the information card may occur. For example, assume that when the area ratio b of the main body image ≤ M < c, it is the first display state, corresponding to a smaller-sized information card (for the convenience of description, it can be called a "small card"). When the area ratio c of the main body image ≤ M < d, it is the second display state, corresponding to a larger-sized information card (for the convenience of description, it can be called a "large card"). If the area ratio of the main body image changes repeatedly near c, for example, assume c = 0.5, then the situation may occur that M changes back and forth between 0.49 and 0.51. At this time, the phenomenon of rapid and repeated switching between the "small card" and the "large card" may occur in the image frame, and this phenomenon can be called the "jitter" phenomenon of the information card.

[0156] To solve the "jitter" phenomenon during the switching of the information card type, in the embodiments of the present application, a method of adding a buffer at the boundary of various main body states is adopted to achieve this. For example, under normal circumstances, the division method of various main body states can be:

[0157] a < M < b is the discovery state;

[0158] b <= M < c is the state of displaying the "small card";

[0159] c <= M < d is the state of displaying the "large card";

[0160] d <= M < e is the interaction state.

[0161] However, to solve the above "jitter" phenomenon during the switching of the information card type, a buffer n can be added at the boundary. At this time, the division method of various main body states can be:

[0162] a < M < b is the discovery state;

[0163] b + n <= M < c is the state of displaying the small card;

[0164] c + n <= M < d is the state of displaying the large card;

[0165] d + n <= M < e is the interaction state.

[0166] This allows for a buffer of size "n" at the boundary between two different subject states, instead of using a simple numerical value as the boundary. In this way, if a user experiences "hand tremor" at the boundary between two states of a subject during image stream acquisition, as long as the change in the subject image area caused by the tremor is within "n", the information card type switch will not be triggered, thus solving the information card "jitter" problem at subject type boundaries. Of course, in practical implementation, other judgment parameters can be used to determine the subject state, such as the size and clarity of the subject image.

[0167] Furthermore, during the specific process of switching information card types, since it involves switching from a first-type information card to a second-type information card—two different information cards—it's essentially replacing the first-type information card with the second-type. Therefore, the switching process can be controlled. For example, in one implementation, the switching process can be divided into two stages: the first stage gradually reduces the transparency of the first-type information card currently projected in the real-time image stream (e.g., the transparency can be gradually reduced from 1 to 0, meaning the first-type information card gradually disappears); the second stage projects the second-type information card into the real-time image stream and gradually increases its transparency (e.g., the transparency can be gradually increased from 0 to 1, meaning the second-type information card gradually appears), and so on.

[0168] In summary, the solution provided in Embodiment 1 of this application allows for the pre-production of real-world merchandise displays in a physical space, followed by 3D reconstruction of the spatial structure of the physical space. The real-world images are then stitched and matched into the 3D-reconstructed virtual space structure to generate a virtual space scene. When a user requests to browse product information in VR mode, the pre-generated virtual space scene can be displayed using VR technology, and subject recognition can be performed on the image frames within the virtual space scene. Then, search results for product information related to the identified subject can be obtained, and information cards to be displayed can be generated based on these search results and mapped to the corresponding position of the subject in the image frame. This approach allows users to experience "online shopping," and the real-world image frames within the virtual space scene can serve as the starting point for "image-based product search." Through subject recognition, product search results are automatically obtained and mapped back to the virtual space scene in the form of information cards, enriching the information within the virtual space scene created through real-world photography and providing users with a novel product search interaction experience.

[0169] Furthermore, to facilitate the creation, updating, and maintenance of virtual space scenes, this application embodiment also provides a robotic device. This device can capture real-world images of merchandise displays in physical spaces to generate virtual space scenes, which can be updated daily. Additionally, the robotic device can provide shopping guidance services to users entering physical spaces. This includes recognizing objects while collecting images of physical merchandise displayed in the mall, generating information cards from search results, and then mapping these cards onto the image stream using AR technology. In this way, a closed loop of "online shopping" and "offline shopping" can be formed using the robotic device, providing users with image-based product search services both online and offline.

[0170] Furthermore, whether "shopping online" through a user terminal device or "shopping offline" through a robot device, the image frames are dynamically changing, causing the position and state of the main image within the image frame to change accordingly. Therefore, in this embodiment, different types of information cards can be provided for different main states, and the information cards can follow the movement of the main body.

[0171] Regarding the issue of discontinuous motion trajectories encountered by information cards during the movement of the main body, a preferred embodiment of this application can solve this problem through segmented fitting. Regarding collisions between information cards and screen edges, two relative positional relationships between the information card and the center point of the main body can be pre-defined; when such collisions occur, the relative positional relationship can be switched to resolve the issue. Regarding collisions between information cards, the priority of different types of information cards can be determined based on the priority relationship between different main body states, thus identifying lower-priority information cards as those that need to be moved. Alternatively, multiple layers can be created, with information cards corresponding to different main body states displayed on different layers.

[0172] Furthermore, in MR glasses scenarios, the interactive method of adding products to the shopping cart by having users "reach out and grab" information cards can be combined to further enhance the fun of the interaction and user participation.

[0173] Example 2

[0174] This second embodiment provides a robotic device that can be applied to physical spaces.

[0175] The robot device is used to provide users entering the physical space with a shopping guide service based on AR mode. The shopping guide service includes: real-time image stream acquisition of the physical merchandise display in the physical space, subject recognition from the real-time image stream, obtaining product information search results related to the identified subject, generating information cards to be displayed based on the product information search results, and displaying them.

[0176] In addition, the robot equipment can also be used to take real-scene photos of the physical merchandise display in the physical space, so as to generate a virtual space scene by stitching and matching the real-scene photos into the three-dimensional reconstructed virtual space structure through three-dimensional reconstruction of the spatial structure of the physical space.

[0177] In the process of VR displaying the virtual space scene through the user's associated terminal device, subject recognition is performed from the real-scene image frame corresponding to the current display perspective, search results for product information related to the identified subject are obtained, information cards to be displayed are generated based on the search results for product information, and the information cards are mapped to the virtual space scene for display.

[0178] Example 3

[0179] It should be noted that the solution provided in this application can also be applied to other physical spaces, such as botanical gardens, zoos, museums, and exhibition halls. Real-world images of the plants, animals, and exhibits can be taken in advance, and corresponding virtual spaces can be generated based on these images. When a user visits this virtual space, subject recognition (the specific subject could be a plant, animal, exhibit, etc.) can be performed from the real-world image frames within the virtual space. Based on the identified subject, a search can be performed to obtain search results for that specific subject, such as encyclopedic information. Then, specific information cards can be generated based on the search results and mapped to the corresponding location of the subject in the virtual space. This allows users to obtain more additional information about the specific subject they see while browsing the virtual space.

[0180] Therefore, this third embodiment also provides an information display method, see [link to example]. Figure 5 The method may include:

[0181] S501: Display a pre-generated virtual space scene based on VR technology. The virtual space scene is generated in the following way: take real-scene photos of the entities in the physical space, reconstruct the spatial structure of the physical space in three dimensions, and stitch the real-scene photos into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0182] S502: Perform subject identification from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene;

[0183] S503: Obtain search results for information related to the identified subject image;

[0184] S504: Generate an information card to be displayed based on the information search results, and map the information card to the virtual space scene for display.

[0185] For the parts not described in detail in Embodiments 2 and 3, please refer to Embodiment 1 and other parts of this application specification, which will not be repeated here.

[0186] It should be noted that the embodiments of this application may involve the use of user data. In practical applications, user-specific personal data may be used in the scheme described herein within the scope permitted by applicable laws and regulations, provided that it complies with the applicable laws and regulations of the country (e.g., with the user's explicit consent, with the user being properly notified, etc.).

[0187] Corresponding to Embodiment 1, this application also provides an apparatus for providing product search result information, see [link to embodiment]. Figure 6 The device may include:

[0188] VR display unit 601 is used to respond to a request to browse product information through virtual reality (VR) mode, and to display a pre-generated virtual space scene based on VR technology. The virtual space scene is generated by: taking real-scene photos of the physical product display in the physical space, reconstructing the spatial structure of the physical space in three dimensions, and stitching the real-scene photos into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0189] The subject recognition unit 602 is used to perform subject recognition from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene;

[0190] The product search result acquisition unit 603 is used to acquire product information search results related to the identified subject;

[0191] The information card display unit 604 is used to generate information cards to be displayed based on the product information search results, and to map the information cards to the virtual space scene for display.

[0192] The real-scene images are obtained by robotic equipment in the physical space taking real-scene photos of the physical merchandise display in the physical space.

[0193] Specifically, the physical space contains multiple different subspaces, and the physical goods include physical goods displayed in the subspaces.

[0194] The real-scene images are obtained by the robotic equipment in the physical space taking real-scene photos of the physical merchandise display in multiple different sub-spaces of the physical space.

[0195] When performing three-dimensional reconstruction of the spatial structure, the spatial distribution structure of multiple subspaces in the physical space is reconstructed in three dimensions.

[0196] In addition, the robotic device can also provide shopping guidance services based on augmented reality (AR) mode to users entering the physical space:

[0197] The shopping guide service provided by the robot device includes: real-time image stream acquisition of the physical merchandise display in the physical space, subject recognition from the real-time image stream, obtaining product information search results related to the identified subject, generating information cards to be displayed based on the product information search results, and displaying them.

[0198] When displaying the information card through the robot device, the information card can be mapped to the location of the subject in the corresponding image frame for display.

[0199] Alternatively, the information card can be displayed using holographic projection.

[0200] In addition to displaying the information card via holographic projection, the 3D model of the product corresponding to the information card can also be displayed via holographic projection.

[0201] Specifically, the subject identification unit can be used for:

[0202] During the process of dynamically changing the viewpoint and / or performing forward and backward operations in the virtual space scene, subject recognition is performed from multiple real-scene captured image frames, and the state changes of the subject in the multiple real-scene captured image frames are obtained. The state is related to the position of the subject in the real-scene captured image frames and / or the area ratio of the subject image on the screen.

[0203] The information card display unit can be specifically used for:

[0204] Information cards to be displayed are generated based on the product information search results and the status of the subject, wherein different subject statuses correspond to different information card types.

[0205] The device may further include:

[0206] The reading unit is used to read the latest position of the identified subject according to a preset period as the information card moves along with the position of the subject; wherein the time interval is related to the persistence of human vision and is greater than the time interval between each frame in the image stream.

[0207] An animation generation unit is used to generate an animation for the currently displayed information card to smoothly move from the current position to the latest position after reading the latest position, and to keep the position of the information card unchanged at the latest position until a new latest position is read in the next cycle, at which time a new animation is generated.

[0208] The relative positional relationship between the information card and the anchor point of the main body is varied; the position of the anchor point is determined based on the center point of the main body.

[0209] The device may further include:

[0210] The relative position relationship switching unit is used to determine whether the information card will collide with the edge of the screen in the current relative position relationship state during the process of the information card moving with the change of the main body position. If so, it switches to other relative position relationships.

[0211] If multiple subjects are identified in the image frame, their respective product information search results are obtained, and information cards of the corresponding type are generated according to their respective states.

[0212] At this point, the device may also include:

[0213] The judgment unit is used to determine whether a collision will occur between the multiple information cards during the process of mapping multiple information cards onto the image frame and moving with the position of their respective subjects.

[0214] The card position moving unit is used to move the display position of the information card corresponding to the lower priority subject according to the different subject status priorities of the information card that collides.

[0215] The information cards corresponding to the different subject states are displayed in different layers, and the layers are arranged in descending order of priority of each subject state.

[0216] Additionally, the device may also include:

[0217] The card type switching unit is used to switch the type of the information card mapped to the virtual space scene if the state of the subject changes during the process of mapping the information card to the location of the corresponding subject in the image frame for display.

[0218] There is a buffer at the boundary of the judgment parameters for different subject states, so that when the judgment parameters change within the buffer, the current subject state and the type of the corresponding information card remain unchanged.

[0219] The state of the subject includes a display state, which includes a first display state and a second display state; wherein, in the display state, the information card is used to display content related to the product search results; the size and content of the information card corresponding to the first display state and the second display state are different.

[0220] The subject's state also includes a discovery state, and the information card in the discovery state is used to display the identification result of the category to which the subject belongs.

[0221] The state of the subject also includes the interaction state;

[0222] The device may further include:

[0223] The interactive unit is used to redirect the user to the product search results list page or the product details page included in the currently displayed information card after the user enters the interactive state and maintains it for a threshold time.

[0224] When displaying the virtual space scene using VR / Mixed Reality (MR) glasses, the device further includes:

[0225] An operation detection unit is used to add the product corresponding to the information card to the user's associated batch checkout product set when the VR / MR glasses detect that the user has performed an operation of grabbing the information card and moving it in the target direction.

[0226] Corresponding to Embodiment 3, this application also provides an information display device, see [link to embodiment]. Figure 7 The device may include:

[0227] VR display unit 701 is used to display a pre-generated virtual space scene based on VR technology. The virtual space scene is generated by: taking real-scene photos of entities in the physical space, reconstructing the spatial structure of the physical space in three dimensions, and stitching the real-scene photos into the three-dimensional reconstructed virtual space structure to generate the virtual space scene.

[0228] The subject recognition unit 702 is used to perform subject recognition from the real-scene image frame corresponding to the current display viewpoint in the virtual space scene;

[0229] Search unit 703 is used to obtain information search results related to the identified subject image;

[0230] The information card display unit 704 is used to generate information cards to be displayed based on the information search results, and to map the information cards to the virtual space scene for display.

[0231] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the method described in any of the foregoing method embodiments.

[0232] And an electronic device, comprising:

[0233] One or more processors; and

[0234] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.

[0235] in, Figure 8 The architecture of an electronic device is illustrated by example. For example, device 800 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, aircraft, etc.

[0236] Reference Figure 8 The device 800 may include one or more of the following components: a processing component 802, a memory 804, a power supply component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0237] Processing component 802 typically controls the overall operation of device 800, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the methods provided in this disclosure. Furthermore, processing component 802 may include one or more modules to facilitate interaction between processing component 802 and other components. For example, processing component 802 may include a multimedia module to facilitate interaction between multimedia component 808 and processing component 802.

[0238] Memory 804 is configured to store various types of data to support the operation of device 800. Examples of this data include instructions for any application or method operating on device 800, contact data, phonebook data, messages, pictures, videos, etc. Memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0239] Power supply component 806 provides power to various components of device 800. Power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 800.

[0240] Multimedia component 808 includes a screen that provides an output interface between device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When device 800 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.

[0241] Audio component 810 is configured to output and / or input audio signals. For example, audio component 810 includes a microphone (MIC) configured to receive external audio signals when device 800 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 804 or transmitted via communication component 816. In some embodiments, audio component 810 also includes a speaker for outputting audio signals.

[0242] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.

[0243] Sensor assembly 814 includes one or more sensors for providing status assessments of various aspects of device 800. For example, sensor assembly 814 may detect the on / off state of device 800, the relative positioning of components such as the display and keypad of device 800, changes in the position of device 800 or a component of device 800, the presence or absence of user contact with device 800, the orientation or acceleration / deceleration of device 800, and temperature changes of device 800. Sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 814 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.

[0244] Communication component 816 is configured to facilitate wired or wireless communication between device 800 and other devices. Device 800 can access wireless networks based on communication standards, such as WiFi, or mobile communication networks such as 2G, 3G, 4G / LTE, and 5G. In one exemplary embodiment, communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0245] In an exemplary embodiment, device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.

[0246] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, which can be executed by a processor 820 of device 800 to perform the method provided by the present disclosure. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0247] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0248] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0249] The method and electronic device for providing product search results information provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and its core ideas. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A method for providing product search result information, characterized in that, include: In response to a request to browse product information through virtual reality (VR) mode, a pre-generated virtual space scene is displayed based on VR technology. The virtual space scene is generated in the following way: real-scene photography of the physical product display in the physical space, three-dimensional reconstruction of the spatial structure of the physical space, and stitching and matching the real-scene photography images into the three-dimensional reconstructed virtual space structure to generate the virtual space scene. During the process of dynamically changing the viewpoint and / or performing forward and backward operations in the virtual space scene, subject recognition is performed from multiple real-scene captured image frames, and the position and state changes of the subject in the multiple real-scene captured image frames are obtained. The subject state is related to the position of the subject in the real-scene captured image frames and / or the area ratio of the subject image on the screen. Obtain search results for product information related to the identified entity; Information cards to be displayed are generated based on the product information search results and the subject status, and the information cards are mapped to the virtual space scene for display. Different subject statuses correspond to different information card types, and the subject statuses include discovery status, display status, and / or interaction status.

2. The method according to claim 1, characterized in that, The real-scene images are obtained by robotic equipment in the physical space taking real-scene photos of the physical merchandise display in the physical space.

3. The method according to claim 2, characterized in that, Also includes: The robotic device provides augmented reality (AR)-based shopping guidance services to users entering the physical space. The shopping guide service provided by the robot device includes: real-time image stream acquisition of the physical merchandise display in the physical space, subject recognition from the real-time image stream, obtaining product information search results related to the identified subject, generating information cards to be displayed based on the product information search results, and displaying them.

4. The method according to claim 3, characterized in that, When displaying the information card through the robotic device, the following is included: The information card is mapped to the location of the subject in the corresponding image frame for display.

5. The method according to claim 4, characterized in that, As the information card moves along with the changing position of the subject, the latest position of the identified subject is read at preset time intervals; wherein, the preset time interval is related to the persistence of vision of the human eye and is greater than the time interval between each frame in the image stream; After the latest position is read, an animation is generated for the currently displayed information card to smoothly move from the current position to the latest position, and the position of the information card is kept unchanged at the latest position until a new latest position is read in the next cycle, at which time a new animation is generated.

6. The method according to claim 4, characterized in that, The relative positional relationship between the information card and the anchor point of the main body can be varied; the position of the anchor point is determined based on the center point of the main body. The method further includes: As the information card moves along with the change in the position of the main body, it is determined whether the information card will collide with the edge of the screen in the current relative position relationship. If so, the relative position relationship is switched to another.

7. The method according to claim 4, characterized in that, If multiple subjects are identified in the image frame, the corresponding product information search results are obtained for each subject, and information cards of the corresponding type are generated according to their respective states. The method further includes: During the process of mapping multiple information cards onto the image frame and moving them according to the changes in the position of their respective subjects, it is determined whether a collision will occur between the multiple information cards. If so, the display position of the information card corresponding to the subject with the lower priority will be moved according to the different subject status priorities of the information card that collides.

8. The method according to claim 4, characterized in that, Also includes: During the process of mapping the information card to the location of the corresponding subject in the image frame for display, if the subject's state changes, the type of information card mapped to the virtual space scene will be switched.

9. The method according to claim 4, characterized in that, The display states include a first display state and a second display state; wherein, in the display state, the information card is used to display content related to the product search results; the size and content of the information cards corresponding to the first display state and the second display state are different.

10. A robotic device, characterized in that, The robotic equipment is used in physical spaces; The robot device is used to provide AR-based shopping guide services to users entering the physical space. The shopping guide service includes: real-time image stream acquisition of the physical merchandise display in the physical space; subject recognition from the real-time image stream; obtaining the position and status changes of the subject in multiple real-time image streams; obtaining product information search results related to the identified subject; generating and displaying information cards based on the product information search results and the subject status. The subject status is related to the position of the subject in the real-scene image frame and / or the area ratio of the subject image on the screen. Different subject statuses correspond to different information card types. The subject status includes discovery status, display status, and / or interactive status.

11. An information display method, characterized in that, include: The virtual space scene is generated by displaying a pre-generated virtual space scene based on VR technology. The virtual space scene is generated by taking real-scene photos of the entities in the physical space and reconstructing the spatial structure of the physical space in three dimensions. The real-scene photos are then stitched and matched into the three-dimensional reconstructed virtual space structure to generate the virtual space scene. During the process of dynamically changing the viewpoint and / or performing forward and backward operations in the virtual space scene, subject recognition is performed from multiple real-scene captured image frames, and the position and state changes of the subject in the multiple real-scene captured image frames are obtained. The subject state is related to the position of the subject in the real-scene captured image frames and / or the area ratio of the subject image on the screen. Based on the identified subject image, obtain search results for information related to the subject; Information cards to be displayed are generated based on the information search results and the subject status, and the information cards are mapped to the virtual space scene for display. Different subject statuses correspond to different information card types, and the subject status includes discovery status, display status, and / or interaction status.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program performs the steps of the method described in any one of claims 1 to 9, 11.

13. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 9, 11.