Virtual reality device and VR scene image recognition method

By detecting the type of content in virtual reality devices and generating appropriate display methods for the recognition results, the problem of inaccurate image recognition in virtual reality devices is solved, thereby improving the accuracy of the recognition results and the user experience.

CN114299407BActive Publication Date: 2026-02-03HISENSE VISUAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202011379185.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-11-30
Publication Date
2026-02-03
Estimated Expiration
2040-11-30

AI Technical Summary

Technical Problem

Traditional virtual reality devices suffer from screen distortion due to optical component distortion, resulting in inaccurate image recognition results. This is especially true for different types of content, where the degree of distortion varies, making it impossible to display the recognition results correctly.

Method used

Virtual reality devices generate recognition results by detecting the source type of the image to be recognized, and display the recognition results in the user interface according to the source type, or by establishing a communication connection with a server, allowing the server to complete the image recognition and display the results.

Benefits of technology

It achieves the correct display of recognition results by using appropriate coordinate mapping methods according to different video source types, solving the problem of inaccurate image recognition in traditional virtual reality devices and improving the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114299407B_ABST
    Figure CN114299407B_ABST
Patent Text Reader

Abstract

The application provides a virtual reality device and a VR scene image recognition method. After obtaining a user input image recognition control instruction, the method can detect the type of a to-be-recognized image, generate a recognition result according to an image recognition algorithm, and display the recognition result in a user interface according to the type of the image. The method can adopt different coordinate mapping modes according to different image sources, so as to correctly display the recognition result in the user interface, and solve the problem that a traditional virtual reality device cannot accurately display a recognition result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of virtual reality equipment technology, and in particular to a virtual reality device and a VR scene image recognition method. Background Technology

[0002] Virtual Reality (VR) technology is a display technology that uses computers to simulate virtual environments, thereby creating a sense of immersion. A VR device is a device that uses virtual display technology to present virtual images to users to achieve an immersive experience. Typically, a VR device includes two display screens to display virtual content, one for each of the user's left and right eyes. When the content displayed on the two screens comes from different perspectives of the same object, it can provide the user with a stereoscopic viewing experience.

[0003] In some applications, image recognition can be performed on the content displayed by virtual reality (VR) devices. For example, image analysis can be used to locate people or specific targets in an image. To perform image recognition, VR devices can take screenshots of the displayed content and execute image recognition programs on the resulting screenshots. However, due to the distortion effects of the optical components used by VR devices, the content displayed on the screen is distorted, deviating significantly from the actual image. Furthermore, the degree of distortion varies depending on the type of content, making it difficult to display the image recognition results correctly. Summary of the Invention

[0004] This application provides a virtual reality device and a VR scene image recognition method to solve the problem that traditional virtual reality devices cannot accurately display recognition results.

[0005] In a first aspect, this application provides a virtual reality device, including: a display and a controller. The display is configured to display a user interface; the controller is configured to perform the following program steps:

[0006] Obtain the control commands input by the user to initiate image recognition;

[0007] In response to the control command, the source type of the image to be identified is detected;

[0008] Generate the recognition result of the image to be recognized;

[0009] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0010] Based on the aforementioned virtual reality device, the first aspect of this application also provides a VR scene image recognition method, applied to a virtual reality device, the method comprising:

[0011] Obtain the control commands input by the user to initiate image recognition;

[0012] In response to the control command, the source type of the image to be identified is detected;

[0013] Generate the recognition result of the image to be recognized;

[0014] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0015] As can be seen from the above technical solutions, the virtual reality device and VR scene image recognition method provided in the first aspect of this application can detect the source type of the image to be recognized after obtaining the user's input image recognition control command, generate a recognition result according to the image recognition algorithm, and display the recognition result in the user interface according to the source type. The method can use different coordinate mapping methods according to different source images, thereby correctly displaying the recognition result in the user interface and solving the problem that traditional virtual reality devices cannot accurately display recognition results.

[0016] Secondly, this application also provides a virtual reality device, including: a display, a communicator, and a controller. The display is configured to display a user interface; the communicator is configured to connect to a server; and the controller is configured to execute the following program steps:

[0017] Obtain the control commands input by the user to initiate image recognition;

[0018] In response to the control command, the source type of the image to be identified is detected;

[0019] The communicator sends an image recognition request to the server.

[0020] Receive the identification results fed back by the server;

[0021] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0022] Based on the aforementioned virtual reality device, a second aspect of this application also provides a VR scene image recognition method, applied to a virtual reality device, the method comprising:

[0023] Obtain the control commands input by the user to initiate image recognition;

[0024] In response to the control command, the source type of the image to be identified is detected;

[0025] The communicator is used to send an image recognition request to the server.

[0026] Receive the identification results fed back by the server;

[0027] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0028] As can be seen from the above technical solutions, the virtual reality device and VR scene image recognition method provided in the second aspect of this application can establish a communication connection between the virtual reality device and the server. After the virtual reality device receives the control commands input by the user and detects the source type of the image to be recognized, it sends an image recognition request to the server. The server can then return the image recognition result according to the request, and the virtual reality device displays the recognition result in the user interface according to the source type of the image to be recognized. This method can delegate the image recognition process to the server, alleviating the processing burden on the virtual reality device and enabling accurate display of the recognition result in the user interface, thus solving the problem that traditional virtual reality devices cannot accurately display recognition results. Attached Figure Description

[0029] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a schematic diagram of the display system structure including a virtual reality device in an embodiment of this application;

[0031] Figure 2 This is a schematic diagram of the global interface of the VR scene in the embodiments of this application;

[0032] Figure 3 This is a schematic diagram of the recommended content area of ​​the global interface in this embodiment of the application;

[0033] Figure 4 This is a schematic diagram of the application shortcut operation entry area of ​​the global interface in this application embodiment;

[0034] Figure 5 This is a schematic diagram of the floating objects on the global interface in an embodiment of this application;

[0035] Figure 6a This is a schematic diagram of a VR screen in an embodiment of this application;

[0036] Figure 6b This is a schematic diagram of the person recognition results in an embodiment of this application;

[0037] Figure 6c This is a schematic diagram of the building identification results in the embodiments of this application;

[0038] Figure 7 This is a flowchart illustrating a VR scene image recognition method according to an embodiment of this application;

[0039] Figure 8 This is a schematic diagram of the initial state of the VR scene in the embodiments of this application;

[0040] Figure 9 This is a schematic diagram showing the image effect in the embodiments of this application;

[0041] Figure 10 This is a schematic diagram showing the effect of the recognition results in the embodiments of this application;

[0042] Figure 11 This is a schematic diagram illustrating the process of generating recognition results based on the source material type in an embodiment of this application;

[0043] Figure 12 This is a schematic diagram of the initial display state of the 3D source material in an embodiment of this application;

[0044] Figure 13 This is a schematic diagram showing the 3D video source recognition results in an embodiment of this application;

[0045] Figure 14 This is a schematic diagram of the initial display state of the 360-degree panoramic video source in an embodiment of this application;

[0046] Figure 15 This is a schematic diagram showing the 360° panoramic image source recognition results in an embodiment of this application;

[0047] Figure 16 This is a schematic diagram of the coordinates of the recognition results in the embodiments of this application;

[0048] Figure 17 This is a schematic diagram of the coordinate mapping state of the recognition result in the embodiments of this application;

[0049] Figure 18 This is a flowchart illustrating another VR scene image recognition method in an embodiment of this application. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the exemplary embodiments of this application clearer, the technical solutions in the exemplary embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described exemplary embodiments are only some embodiments of this application, and not all embodiments.

[0051] Based on the exemplary embodiments shown in this application, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this application. Furthermore, although the disclosures in this application are presented by way of one or more exemplary examples, it should be understood that each aspect of these disclosures can constitute a complete technical solution on its own.

[0052] It should be understood that the terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such use of data can be interchanged where appropriate, for example, to allow implementation in orders other than those given in the embodiments illustrated or described in this application.

[0053] Furthermore, the terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclusively include, for example, a product or device that includes a series of components is not necessarily limited to those that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.

[0054] As used in this application, the term "module" means any known or subsequently developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with that element.

[0055] Throughout this specification, references to "multiple embodiments," "some embodiments," "one embodiment," or simply "embodiment" indicate that a specific feature, structure, or characteristic described in connection with that embodiment is included in at least one embodiment. Therefore, phrases such as "in multiple embodiments," "in some embodiments," "in at least another embodiment," or "in an embodiment" appearing throughout this specification do not necessarily refer to the same embodiment. Furthermore, in one or more embodiments, specific features, structures, or characteristics can be combined in any suitable manner. Therefore, without limitation, a specific feature, structure, or characteristic shown or described in connection with one embodiment may be combined, in whole or in part, with features, structures, or characteristics of one or more other embodiments. Such modifications and variations are intended to be included within the scope of this application.

[0056] In this embodiment, the virtual reality device 500 generally refers to a display device that can be worn on a user's face to provide an immersive experience, including but not limited to VR glasses, augmented reality (AR) devices, VR gaming devices, mobile computing devices, and other wearable computers. The virtual reality device 500 can operate independently or be connected to other smart display devices as an external device, such as smart TVs, computers, tablets, servers, etc.

[0057] The virtual reality device 500, when worn on a user's face, displays media assets, providing close-up images to the user's eyes for an immersive experience. To present the media assets, the virtual reality device 500 can include multiple components for displaying the images and for wearing on the face. Taking VR glasses as an example, the virtual reality device 500 can include a shell, temples, an optical system, a display component, a posture detection circuit, and an interface circuit. In practical applications, the optical system, display component, posture detection circuit, and interface circuit can be housed within the shell to display the specific images; the temples are connected to both sides of the shell for wearing on the user's face.

[0058] When in use, the attitude detection circuit has built-in attitude detection components such as gravity acceleration sensor and gyroscope. When the user's head moves or turns, the user's attitude can be detected and the detected attitude data can be transmitted to the controller and other processing components, so that the processing components can adjust the specific screen content in the display component according to the detected attitude data.

[0059] It should be noted that the specific content presented varies depending on the type of virtual reality device 500. For example, as... Figure 1 As shown, for some thin and light VR glasses, the built-in controller generally does not directly participate in the control process of the displayed content. Instead, it sends the posture data to an external device, such as a computer, for processing. The external device then determines the specific screen content to be displayed and sends it back to the VR glasses to display the final screen.

[0060] In some embodiments, the virtual reality device 500 can be connected to the display device 200 and a network-based display system can be built between it and the server 400. Data interaction can be performed in real time between the virtual reality device 500, the display device 200 and the server 400. For example, the display device 200 can obtain media data from the server 400 and play it, and transmit specific screen content to the virtual reality device 500 for display.

[0061] The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. The specific type, size, and resolution of the display device are not limited. Those skilled in the art will understand that the display device 200 can be modified in terms of performance and configuration as needed. The display device 200 can provide broadcast television reception functionality and may also include, but is not limited to, intelligent network television functionality with computer support, including but not limited to, internet television, smart television, and Internet Protocol television (IPTV).

[0062] Display device 200 and virtual reality device 500 also communicate with server 400 via various communication methods. Display device 200 and virtual reality device 500 can communicate via local area network (LAN), wireless local area network (WLAN), and other networks. Server 400 can provide display device 200 with various content and interactive features. For example, display device 200 can interact by sending and receiving information, as well as electronic program guides (EPGs), receiving software updates, or accessing remotely stored digital media libraries. Server 400 can be a cluster or multiple clusters, and may include one or more types of servers. Other network services such as video-on-demand and advertising services can be provided through server 400.

[0063] During data interaction, the user can operate the display device 200 through the mobile terminal 100A and the remote control 100B. The mobile terminal 100A and the remote control 100B can communicate with the display device 200 via a direct wireless connection or a non-direct connection. Specifically, in some embodiments, the mobile terminal 100A and the remote control 100B can communicate with the display device 200 via direct connection methods such as Bluetooth or infrared. When sending control commands, the mobile terminal 100A and the remote control 100B can directly transmit the control command data to the display device 200 via Bluetooth or infrared.

[0064] In other embodiments, the mobile terminal 100A and the remote controller 100B can also access the same wireless network as the display device 200 via a wireless router to establish a non-direct connection communication with the display device 200 through the wireless network. When sending control commands, the mobile terminal 100A and the remote controller 100B can first send the control command data to the wireless router, and then the wireless router forwards the control command data to the display device 200.

[0065] In some embodiments, users can also use the mobile terminal 100A and the remote control 100B to interact directly with the virtual reality device 500. For example, the mobile terminal 100A and the remote control 100B can be used as controllers in a virtual reality scene to achieve functions such as motion-sensing interaction.

[0066] In some embodiments, the display component of the virtual reality device 500 includes a display screen and driving circuitry associated with the display screen. To present a concrete image and provide a stereoscopic effect, the display component may include two display screens, corresponding to the user's left and right eyes respectively. When presenting a 3D effect, the content displayed on the left and right screens will be slightly different, and can respectively display the left and right cameras used during the filming of the 3D source material. Because the user observes the image content with their left and right eyes, a more stereoscopic display image can be observed when the device is worn.

[0067] The optical system in the virtual reality device 500 is an optical module composed of multiple lenses. Positioned between the user's eyes and the display screen, the optical system increases the optical path through the refraction of light signals by the lenses and the polarization effect of polarizers on the lenses, ensuring that the content displayed is clearly presented within the user's field of vision. Furthermore, to accommodate different users' visual acuity, the optical system also supports focusing. This involves adjusting the position of one or more lenses using a focusing component, changing the distance between the lenses, and thus altering the optical path and adjusting the image clarity.

[0068] The interface circuit of the virtual reality device 500 can be used to transmit interactive data. Besides transmitting posture data and display content data, in practical applications, the virtual reality device 500 can also connect to other display devices or peripherals through the interface circuit to achieve more complex functions through data interaction with the connected devices. For example, the virtual reality device 500 can connect to a display device through the interface circuit to output the displayed image to the display device in real time. As another example, the virtual reality device 500 can also connect to a controller through the interface circuit, which can be held and operated by the user to perform related operations in the VR user interface.

[0069] The VR user interface can be presented in various different UI layouts based on user operations. For example, the user interface may include a global interface, such as the global UI after the AR / VR terminal is started. Figure 2 As shown, the global UI can be displayed on the display screen of the AR / VR terminal or on the display device's monitor. The global UI may include a recommended content area 1, a business category extension area 2, an application quick operation entry area 3, and a floating element area 4.

[0070] Recommended content area 1 is used to configure different category tabs; within these tabs, media assets, special features, etc., can be configured; the media assets may include 2D films, educational courses, tourism, 3D, 360-degree panoramic, live streaming, 4K films, applications, games, and other businesses with media asset content, and these tabs can select different template styles and support simultaneous recommendation and arrangement of media assets and special features, such as... Figure 3 As shown.

[0071] The Business Category Extension Area 2 supports configuring extended categories for different business types. When a new business type is added, a separate tab can be configured to display the corresponding page content. The extended categories in Business Category Extension Area 2 can also be sorted, and businesses can be removed from the platform. In some embodiments, Business Category Extension Area 2 may include the following content: Movies & TV, Education, Travel, Applications, and My Account. In some embodiments, Business Category Extension Area 2 is configured to display large business category tabs and supports configuring more categories; its icons are configurable, such as... Figure 3 As shown.

[0072] The application quick access area 3 can prioritize pre-installed applications for promotional purposes, and supports configuring special icon styles to replace default icons. Multiple pre-installed applications can be specified. In some embodiments, the application quick access area 3 also includes left-hand and right-hand movement controls for selecting different icons, such as... Figure 4 As shown.

[0073] The floating area 4 can be configured to be located above the left or right diagonal side of a fixed area, and can be configured as a replaceable image or a jump link. For example, after receiving a confirmation operation, the floating area can jump to an application or display a specified function page, such as... Figure 5 As shown. In some embodiments, the suspended object may not be configured with a jump link and may be used simply for visual display.

[0074] In some embodiments, the global UI also includes a status bar at the top for displaying the time, network connection status, battery status, and more quick access points. When an icon is selected using the AR / VR terminal's controller, the icon will display a text prompt that expands to the left and right, and the selected icon will stretch and expand to the left and right according to its position.

[0075] For example, after selecting the search icon, the search icon will display the text "Search" and the original icon. Clicking the icon or text will take you to the search page. Another example is that clicking the favorites icon will take you to the favorites tab, clicking the history icon will display the history page by default, clicking the search icon will take you to the global search page, and clicking the message icon will take you to the message page.

[0076] In some embodiments, interaction can be performed through peripheral devices, such as the controller of an AR / VR terminal, which can operate the user interface of the AR / VR terminal, including a back button; a home button, which can be long-pressed to achieve a reset function; volume up and down buttons; and a touch area that can realize the functions of clicking, sliding, holding and dragging the focus.

[0077] Users can access different scene interfaces through the global interface, for example... Figure 6aAs shown, users can access the browsing interface through the "Browse Interface" entry in the global interface, or by selecting any media asset in the global interface. Within the browsing interface, the virtual reality device 500 can create 3D scenes using the Unity 3D engine and render specific visual content within those scenes.

[0078] Within the browsing interface, users can view specific media assets. To enhance the viewing experience, different virtual scene controls can be set up within the interface to present specific scenes or enable real-time interaction in conjunction with the media assets. For example, panels within a Unity 3D scene can be set up in the browsing interface to display image content, which, combined with other virtual home controls, can create a cinema screen effect.

[0079] The virtual reality device 500 can display operation UI content in the browsing interface. For example, a list UI can be displayed in front of the display panel in the Unity 3D scene. The list UI can display icons of media assets currently stored locally on the virtual reality device 500, or icons of network media assets that can be played on the virtual reality device 500. Users can select any icon in the list UI, and the selected media asset will be displayed in real time in the display panel.

[0080] While displaying specific images of media assets, the virtual reality device 500 can also perform image recognition on the displayed content, identifying specific images from the displayed image and marking them. For example, it can identify targets such as people, buildings, and key landmarks in the displayed image and mark the target locations. The virtual reality device 500 displays target markings simultaneously with the image, such as using a bounding box to select identified people.

[0081] Media assets that can be displayed in a Unity 3D scene can be in various forms such as images and videos. Furthermore, due to the display characteristics of VR scenes, media assets displayed in a Unity 3D scene must include at least 2D images or videos, 3D images or videos, and 360-degree panoramic images or videos.

[0082] Among them, 2D images or videos are traditional image or video files that can be displayed on two display screens of the virtual reality device 500 when displayed. In this application, 2D images or videos are collectively referred to as 2D source material. 3D images or videos, i.e., 3D source material, are images or videos created by at least two cameras capturing the same object from different angles. They can be displayed on two display screens of the virtual reality device 500. 360-degree panoramic images or videos, i.e. 360-degree panoramic source material, are 360-degree panoramic images obtained through a panoramic camera or special shooting methods. They can be displayed by creating a display sphere in a Unity 3D scene.

[0083] Because the types of video sources displayed differ, the display effects of the recognition results will vary depending on the type of video source. For example, for 2D images or videos, the recognition frame of the result can be directly displayed on the display panel. However, for 360-degree panoramic video sources, which need to be displayed on a spherical surface, and since the recognition frame cannot be directly displayed on a spherical surface, recognition indicator points can be used to mark the position of the recognition result.

[0084] It should be noted that the recognition results can also be marked in other ways, such as with indicator lines, circles, ovals, triangles, rhombuses, or other geometric shapes, or with display effects such as highlighting or color changes. Furthermore, while displaying the recognition results, explanatory text can be added to explain the results. For example... Figure 6b As shown, when recognizing a person's image, information such as the gender and age of the recognized person can be displayed near the recognition box; for example... Figure 6c As shown, when a building target is identified, information such as the name of the identified building can be displayed near the identification box to improve the user's actual viewing experience.

[0085] However, for different types of video sources, the images displayed on the left and right screens are different during the display process, or the forms of expression in the Unity 3D scene are different, which will cause distortion or differences between the display result and the original image, causing the recognition result to be misaligned on the display screen and reducing the user experience.

[0086] To accurately display the image recognition results, such as Figure 7 As shown, some embodiments of this application provide a VR scene image recognition method, which can be applied to a virtual reality device 500. The method includes the following:

[0087] The user inputs a control command to the virtual reality device 500 to initiate image recognition. Upon receiving the command, the device recognizes the image and displays the recognition result. This image recognition result display serves as an auxiliary display function when the virtual reality device 500 displays media assets. Therefore, the user can choose whether to enable the real-time display of recognition results as needed. For example, the user can enable the "AI" function in the settings interface, which will allow the virtual reality device 500 to perform real-time image recognition while displaying media asset content, and then display the recognition result within the media asset content.

[0088] like Figure 8 , Figure 9As shown, when the assisted display function is enabled, opening any media resource and entering the browsing interface signifies that the user has input a control command to initiate image recognition. This control command can be input by the user using a remote control or motion-sensing controller to move the focus cursor to any image icon in the user interface and then clicking the confirmation or play button. Conversely, when the assisted display function is disabled, selecting the toggle button in the browsing interface and clicking the confirmation button to enable the assisted display function signifies that the user has input a control command to initiate image recognition. Control commands can also be input in other ways, such as through a voice system or an external smart terminal.

[0089] After receiving the control commands input by the user, the virtual reality device 500 can begin image recognition according to the control commands. Since the virtual reality device 500 performs image recognition and displays the image recognition results in different ways depending on the type of source material displayed, the source material type of the image to be recognized can be detected before image recognition is performed. The source material type includes at least 2D source material, 3D source material, and 360-degree panoramic source material.

[0090] To detect the source type of media, the controller can extract information such as the media asset category, format, extension, and file description after receiving a control command, thereby determining the source type of the currently displayed media asset. For example, for network resources presented in the user interface, the source type of the media asset can be indicated in the file description while sharing the media asset.

[0091] The source type of the currently displayed media asset can also be determined by combining the specific image content. For example, if the image file of the displayed media asset has the extension ".jpg", and by analyzing the similarity between the left and right sides of the image, if the similarity between the two sides is small, the source type of the current image to be identified can be determined to be a 2D source; if the similarity between the two sides is large, the source type of the current image to be identified can be determined to be a 3D source.

[0092] After detecting the source type of the displayed media assets, the controller can perform image recognition on the image to be recognized according to the specific recognition method for that type of image, in order to generate a recognition result for the image to be recognized. The specific image recognition method is not limited in this embodiment. For example, image recognition can employ a recognition model, whereby the image to be recognized can be input into the recognition model, and the recognition model outputs the recognition result.

[0093] Different recognition methods can be selected based on specific user needs and application scenarios to obtain different recognition results. When processing different media asset files, different types of recognition models can be used. After detecting the source type of the image to be recognized, the image to be recognized can be input into the recognition model according to the input method corresponding to that source type. The recognition model can calculate the image to be recognized using a preset image recognition algorithm to obtain the recognition result.

[0094] For example, when using the virtual reality device 500 to simulate travel, a scene recognition model can be built into the application. Users wearing the virtual reality device 500 can browse different scenes, and at the same time, the image recognition algorithm can identify specific targets in the scene, thereby marking the location of the scenic spot, its name, explanation and other relevant information.

[0095] After generating the recognition results, the virtual reality device 500 can display them in the user interface. Recognition results for different video source types can be displayed in different ways. For example, ... Figure 10 As shown, for 2D or 3D source images to be identified, the image can be displayed in the display panel of the Unity 3D scene, and a recognition bounding box can be displayed on the image to select the target. For 360-degree panoramic source images, recognition markers can be located on the display sphere of the Unity 3D scene, and these markers can be marked and displayed using guide lines.

[0096] As can be seen from the above technical solutions, the VR scene image recognition method provided in the above embodiments can detect the source type of the image to be recognized after obtaining the user's input image recognition control command, generate recognition results according to the image recognition algorithm, and display the recognition results in the user interface according to the source type. The method can use different coordinate mapping methods according to different source images, thereby correctly displaying the recognition results in the user interface and solving the problem that traditional virtual reality devices 500 cannot accurately display recognition results.

[0097] Because different media sources differ in their visual presentation, the methods for image recognition also differ. For example, 2D images, presented as a single image, can be directly recognized by the recognition model. However, 3D images, presented as two side-by-side images taken from different angles, have slightly different content depending on the relative positions of the cameras. When performing image recognition on 3D images, inputting the entire original image into the recognition model will result in errors due to interference from the two side images. Therefore, if... Figure 11As shown in some embodiments of this application, in order to obtain image recognition results, the step of generating recognition results for the image to be recognized further includes:

[0098] If the source image of the image to be identified is of the first type, the original source image is extracted as the image to be identified;

[0099] Image recognition is performed on the original source image to generate recognition results;

[0100] If the source image of the image to be identified is of the second type, extract the half-side image corresponding to the left or right display in the source image as the image to be identified;

[0101] Image recognition is performed on half of the source image to generate recognition results.

[0102] Before recognizing the image to be recognized, preprocessing can be performed based on the source type of the image. In this embodiment, the source type can include a first type of source and a second type of source. The first type refers to a source type whose content includes only a single image, including but not limited to 2D and 360° panoramic sources; the second type refers to a source type whose content includes two or more images, including but not limited to 3D sources. When the source type of the image to be recognized is detected to be a first type, such as 2D or 360° panoramic, the original image can be directly input into the recognition model for processing to generate the recognition result. Figure 12 , Figure 13 As shown, when the source type of the image to be identified is detected to be a second type such as 3D source, the image to be identified can be cropped and separated to extract the half-side image corresponding to the left or right display in the source image and input it into the recognition model for recognition to generate recognition results.

[0103] For example, in 2D image playback mode, the original 2D image to be displayed can be acquired and shown on a designated panel in the Unity 3D scene. Simultaneously, the Android layer inputs the original image into the recognition model via a recognition request for image recognition. The Android layer is a system layer used to transfer data and instructions between different software layers. In virtual reality devices, layers parallel to the Android layer can also include the application layer and the framework layer. The application layer is configured to present specific algorithms and directly display screen content. The recognition model can be integrated into the application layer, interacting with the system layer through data exchange—that is, acquiring images from the system layer, performing recognition, and simultaneously feeding back the recognition results to the system layer. In 3D image playback mode, after acquiring the original image to be displayed, the left and right images can be displayed on designated panels in the Unity 3D scene. Simultaneously, the Android layer inputs the left half of the original image into the recognition model via a recognition request for image recognition.

[0104] It should be noted that different image or video source types may require different preprocessing methods depending on their image content structure. For example, some 3D source images are arranged horizontally, meaning a frame consists of two halves: the left half displays content on the left side of the screen, and the right half displays content on the right side. In this case, either the left or right half of the source image can be extracted as the image to be identified. Conversely, some 3D source images are arranged vertically, meaning a frame consists of two halves: the upper half displays content on the left side of the screen, and the lower half displays content on the right side. In this case, either the upper or lower half of the source image can be extracted as the image to be identified.

[0105] Furthermore, some 3D content sources use a hybrid image arrangement, meaning that a single frame does not have fixed divisions between regions. Instead, the content displayed on the left and right displays is mixed and arranged. For example, in two adjacent columns of pixels, one column represents the content displayed on the left display, and the other column represents the content displayed on the right display. Multiple columns of pixels are arranged alternately to form a single frame. For 3D content images with a hybrid arrangement, pixel recombination can be performed before sending them for image recognition to separate the content displayed on the left and right displays, obtaining a left image and a right image, one of which can then be used as the image to be recognized.

[0106] As can be seen, in this embodiment, by performing different preprocessing on images of different source types, the images input into the recognition model can retain the specific image content while mitigating the interference of the left and right image content, thereby generating correct recognition results.

[0107] Because the image representation of the images to be identified differs depending on the type of source material, the specific recognition algorithms used for image recognition also vary. For example, for 360-degree panoramic sources, due to the perspective connection during shooting or compositing, the entire 360-degree field of view is displayed in the same image. However, distortion occurs at the bottom of the image during compositing. Therefore, image recognition algorithms using 2D images will be affected by interference from the distorted areas, impacting the recognition results. Thus, in some embodiments, different recognition models can be called according to different source material types. The step of generating the recognition result for the image to be identified further includes:

[0108] The recognition model is invoked according to the source type of the image to be recognized;

[0109] The image to be identified is input into the called recognition model;

[0110] Obtain the recognition result output by the recognition model.

[0111] The recognition model can be pre-built according to different video source types. This application does not limit the specific model building method; it can be obtained through model training or by building an image analyzer. The constructed recognition model can be stored in the memory of the virtual reality device 500 or the display device performing image recognition processing, for use by the controller.

[0112] The controller can call a recognition model according to the source type of the image to be recognized, and input the cropped image to be recognized in the above embodiment into the called recognition model, which then performs recognition processing on the image. After processing the image, the recognition model can output the recognition result, i.e., the controller obtains the recognition result output by the recognition model. Since different recognition models are built for different source types, the recognition model can adapt to the source type of the current image to be recognized, thus obtaining a more accurate recognition result.

[0113] Furthermore, different recognition models can be invoked based on different application scenarios to obtain different recognition results. For example, after receiving control commands input by the user, the controller can also determine the current application scenario to identify the required recognition model group. The recognition model group can include at least three recognition models that can meet the functions of the current scenario, respectively used for image recognition of 2D, 3D, and 360-degree panoramic images. Then, based on the type of image source, a suitable recognition model is selected from the recognition model group.

[0114] Different recognition models will produce different recognition results. For example, for a recognition model trained on a given model, the input recognition result is the classification probability of each region in the image for a specific category.

[0115] In some embodiments, the recognition result may include a result marker and the position of the result marker relative to the image to be recognized; for 2D and 3D source images to be recognized, the result marker is a recognition box, and the position of the result marker includes the coordinates of the upper left and lower right corners of the recognition box; such as Figure 14 , Figure 15 As shown, for a 360-degree panoramic image to be identified, the result is marked as an identification indicator point, and the position of the result is the coordinate of the identification indicator point.

[0116] Because the recognition results for different image source types are presented differently, their final display will also differ. For example, the recognition box needs to be displayed on a flat surface, while the recognition indicator point can be displayed on a curved surface. Therefore, in some embodiments of this application, in order to display the recognition results in the user interface according to the image source type of the image to be recognized, the step of displaying the recognition results further includes:

[0117] The result display area is set in the user interface according to the source type of the image to be identified;

[0118] Extract the coordinate parameters of the result display area in the user interface;

[0119] Coordinate mapping is performed based on the coordinate parameters to display the recognition results in the result display area.

[0120] After generating the recognition results, a result display area can be set up in the Unity 3D scene. The specific form of the display area can be set according to the user interface and virtual reality functions. For example, for a virtual cinema, the display area would be the screen in the virtual cinema. After setting up the result display area, the image to be recognized can also be displayed in the result display area. Obviously, when the image to be recognized is an image from a video, the image displayed in the result display area will also change dynamically.

[0121] Different types of video sources require different display areas. For example, if the video source of the image to be identified is a 2D or 3D video source, a display panel is created in the user interface to display the image to be identified in a tiled manner. If the video source of the image to be identified is a 360-degree panoramic video source, a display sphere is created in the user interface to display the image to be identified around the viewer.

[0122] Since the size and position of the results display area are set according to the specific VR scene, the image to be recognized will be scaled and adjusted according to the size and position of the results display area when it is displayed. The corresponding recognition results also need to be adjusted accordingly when displayed. That is, after setting the results display area, the controller can extract the coordinate parameters of the results display area in the Unity 3D scene and perform coordinate mapping transformations based on these parameters to display the recognition results within the results display area.

[0123] The coordinate parameters include spatial location and region shape data. Specifically, the step of performing coordinate mapping based on the coordinate parameters further includes:

[0124] If the source type of the image to be identified is a 2D source in the first type or a 3D source in the second type, extract the identification mark position from the identification result;

[0125] Obtain the spatial location of the result display area;

[0126] Based on the location of the identification mark and the spatial location, calculate the coordinates of the top left and top right corners of the identification mark in the user interface.

[0127] After generating the image recognition result, the controller can also determine the data type to be extracted based on the source type of the image to be recognized. If the source type of the current image to be recognized is a 2D or 3D source, the recognition marker is a recognition box, meaning the recognition result can be marked using a recognition box. Then, the positions of the recognition markers in the recognition result can be extracted, and the spatial position of the result display area in the Unity 3D scene can be obtained. The spatial position includes the coordinates of the upper left and upper right corners of the result display area.

[0128] After obtaining the spatial location, the coordinates of the top left and top right corners of the identification mark in the user interface are calculated based on the location of the identification mark and the spatial location. Then, the identification box is rendered based on the calculated coordinates of the top left and top right corners of the identification mark and displayed in the results display area.

[0129] For example, such as Figure 16 As shown, the recognition result information includes type: building, location: (x:0.2215, y:0.3325, w:0.5825, h:0495), where x is the x-axis coordinate of the upper left corner of the recognition box / the width W of the original image, y is the y-axis coordinate of the upper right corner of the recognition box / the height H of the original image, w is the width of the recognition box / the width W of the original image, and h is the height of the recognition box / the height H of the original image.

[0130] like Figure 17As shown, the coordinates of the top left corner of the panel in the scene are (LTPx, LTPy, LTPz) and the coordinates of the bottom right corner in the scene are (RBPx, RBPy, RBPz). The coordinates of the recognition box in the recognition result are (x, y, w, h). The coordinates of the top left corner of the recognition box displayed in the scene are (RLx, RLy, RLz) and the coordinates of the bottom right corner are (RRx, RRy, RRz). The coordinate mapping method is to calculate the coordinates of the recognition box in the Unity 3D scene.

[0131] That is, the coordinates of the top left corner of the recognition box are:

[0132] RLx = LTPx + (RBPx - LTPx) * x;

[0133] RLy = LTPy + (RBPy - LTPy) * y;

[0134] RLz = LTPz + (RBPz - LTPz) * x;

[0135] The coordinates of the bottom right corner of the recognition box are:

[0136] RRx=LTPx+(RBPx-LTPx)*(x+w);

[0137] RRy=LTPy+(RBPy-LTPy)*(y+h);

[0138] RRz=LTPz+(RBPz-LTPz)*(x+w);

[0139] As can be seen, by using the above coordinate mapping calculation method, the image recognition results of 2D or 3D video sources can be displayed in the result display area, so that the recognition results can be correctly displayed in the VR scene.

[0140] If the source type of the image to be identified is a 360-degree panoramic source in the first type, extract the identification marker position from the identification result;

[0141] Convert the location of the identification marker into latitude and longitude;

[0142] Obtain the region shape data of the result display area;

[0143] Based on the latitude and longitude and the shape data of the region, calculate the position coordinates of the identification mark in the user interface.

[0144] Since 360° panoramic video sources need to be displayed on a spherical surface, to achieve better display results, when the source type of the image to be recognized is 360° panoramic, the recognition result should be able to be marked on the spherical surface. Therefore, the recognition bounding boxes in the 2D image need to be converted into marker points that can be displayed on the spherical surface.

[0145] When displaying the recognition results, the location of the recognition mark can be extracted from the recognition results first, and the location of the recognition mark can be converted into latitude and longitude information on the display sphere. Then, the radius of the display sphere corresponding to the result display area can be obtained, and the position coordinates of the recognition mark in the user interface can be calculated based on the latitude and longitude and the shape data of the area.

[0146] For example, if the coordinates of the recognition box in the recognition result are (x, y, w, h), and the coordinates of the transformed marker point are (RLx, Rly, RLz), then using the coordinates of the upper left corner of the recognition box as a reference, the recognition box is mapped onto the display sphere. Therefore, the latitude and longitude information can be calculated based on the coordinates of the recognition box and the marker point.

[0147] Wd (longitude) = (x + 90) * π / 180;

[0148] Jd (latitude) = y * π / 180;

[0149] The coordinates of the marked point (RLx, Rly, RLz) are:

[0150] RLx = -r*cos(jd)*cos(wd);

[0151] RLy = -r*sin(jd);

[0152] RLz = r * cos(jd) * sin(wd);

[0153] Where r is the radius of the display sphere, which can be set according to the actual distance in the scene. As can be seen, in the above embodiment, marker points can be used instead of recognition boxes to display the recognition results, thus adapting to the display format of the sphere and enabling image recognition results of 360-degree panoramic source material types to be displayed in VR scenes.

[0154] It should be noted that in the above embodiments, the source material types are described using 2D source material, 3D source material, and 360-degree panoramic source material as examples. Other image recognition methods for source material types that can be conceived by those skilled in the art in conjunction with the source material types in the above embodiments without creative effort are also within the protection scope of this application.

[0155] Based on the above-described VR scene image recognition method, some embodiments of this application also provide a virtual reality device 500, including: a display and a controller, wherein the display is configured to display a user interface; and the controller is configured to execute the following program steps:

[0156] Obtain the control commands input by the user to initiate image recognition;

[0157] In response to the control command, the source type of the image to be identified is detected;

[0158] Generate the recognition result of the image to be recognized;

[0159] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0160] As can be seen from the above technical solutions, the virtual reality device 500 provided in the above embodiments can detect the source type of the image to be recognized after obtaining the user's input image recognition control command, generate a recognition result according to the image recognition algorithm, and display the recognition result in the user interface according to the source type. The virtual reality device 500 can use different coordinate mapping methods according to different source images, thereby correctly displaying the recognition result in the user interface and solving the problem that traditional virtual reality devices 500 cannot accurately display recognition results.

[0161] In the above embodiments, image recognition is performed by the virtual reality device 500. Since the virtual reality device 500 has limited computing and storage capabilities, the image recognition process can also be delegated to other devices. That is, in some embodiments of this application, a VR scene image recognition method is also provided, applied to the virtual reality device 500. The virtual reality device 500 includes a display, a communicator, and a controller, wherein the display is configured to display a user interface; the communicator is configured to connect to a server; and so on. Figure 18 As shown, the method includes the following steps:

[0162] Obtain the control commands input by the user to initiate image recognition;

[0163] In response to the control command, the source type of the image to be identified is detected, including 2D source, 3D source and 360 panoramic source;

[0164] The communicator sends an image recognition request to the server.

[0165] Receive the identification results fed back by the server;

[0166] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0167] The difference between this embodiment and the above embodiment is that, after detecting the source type of the image to be identified, this embodiment can send an image recognition request to the server through a communicator. After receiving the image recognition request, the server can send the image recognition result back to the virtual reality device 500.

[0168] In order for the server to provide image recognition results for the image to be recognized, the image recognition request sent by the virtual reality device 500 should include the image to be recognized. In some embodiments, the virtual reality device 500 may send different image recognition requests depending on the source type of the image to be recognized. For example, for 2D or 360-degree panoramic sources, the image recognition request may include the original source image of the image to be recognized; while for 3D sources, the image recognition request may include the left half of the original source image.

[0169] In this embodiment, the image to be recognized is sent to the server for image recognition, which can reduce the data processing load of the virtual reality device 500 and eliminate the need for the virtual reality device 500 to maintain multiple recognition models, thereby reducing the configuration requirements of the virtual reality device 500.

[0170] Based on the above-described VR scene image recognition method, some embodiments of this application also provide a virtual reality device 500, including: a display, a communicator, and a controller, wherein the display is configured to display a user interface; the communicator is configured to connect to a server; and the controller is configured to execute the following program steps:

[0171] Obtain the control commands input by the user to initiate image recognition;

[0172] In response to the control command, the source type of the image to be identified is detected, including 2D source, 3D source and 360 panoramic source;

[0173] The communicator sends an image recognition request to the server.

[0174] Receive the identification results fed back by the server;

[0175] The recognition result is displayed in the user interface according to the source type of the image to be recognized.

[0176] As can be seen from the above technical solutions, the virtual reality device 500 provided in the above embodiments can establish a communication connection between the virtual reality device 500 and the server. After the virtual reality device 500 obtains the control commands input by the user and detects the source type of the image to be recognized, it sends an image recognition request to the server. The server can then return the image recognition result according to the request, and the virtual reality device 500 displays the recognition result in the user interface according to the source type of the image to be recognized. The virtual reality device 500 can delegate the image recognition process to the server, alleviating the processing burden on the virtual reality device 500 and enabling the correct display of the recognition result in the user interface, thus solving the problem that traditional virtual reality devices cannot accurately display recognition results.

[0177] Similar parts between the embodiments provided in this application can be referred to mutually. The specific implementation methods provided above are only a few examples under the overall concept of this application and do not constitute a limitation on the scope of protection of this application. For those skilled in the art, any other implementation methods extended from the solution of this application without creative effort shall fall within the scope of protection of this application.

Claims

1. A virtual reality device, characterized in that, include: The monitor is configured to display the user interface; The controller is configured as follows: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the source type of the image to be identified is detected; If the source image of the image to be identified is of the first type, the original source image is extracted as the image to be identified; Image recognition is performed on the original source image to generate recognition results; If the source image of the image to be identified is of the second type, extract the half-side image corresponding to the left or right display in the source image as the image to be identified; Image recognition is performed on a half-side image of the source image to generate a recognition result; The recognition result is displayed in the user interface according to the source type of the image to be recognized.

2. The virtual reality device according to claim 1, characterized in that, In the step of generating the recognition result of the image to be recognized, the controller is further configured to: The recognition model is invoked according to the source type of the image to be recognized; The image to be identified is input into the recognition model; Obtain the recognition result output by the recognition model.

3. The virtual reality device according to claim 1 or 2, characterized in that, The recognition result includes a result label and the position of the result label relative to the image to be recognized; For different source material types, the result marker is one or more of the following combinations: recognition box, recognition indicator point, highlight mark, and color change mark; the position of the result marker is a specified point in the result marker area, including the coordinates of the graphic vertex, the graphic midpoint, and the indicator point.

4. A virtual reality device, characterized in that, include: The monitor is configured to display the user interface; The controller is configured as follows: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the media asset information of the image to be identified is extracted, and the source type of the image to be identified is determined based on the media asset information; Generate the recognition result of the image to be recognized; The result display area is set in the user interface according to the source type of the image to be identified; If the source type of the image to be identified is a 2D source in the first type or a 3D source in the second type, the result display area created in the user interface is a display panel. If the source type of the image to be identified is a 360-degree panoramic source in the first type, the result display area created in the user interface is a spherical display area; Extract the coordinate parameters of the result display area in the user interface, the coordinate parameters including spatial location and area shape data; Coordinate mapping is performed based on the coordinate parameters to display the recognition results in the result display area.

5. A virtual reality device, characterized in that, include: The monitor is configured to display the user interface; The controller is configured as follows: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the media asset information of the image to be identified is extracted, and the source type of the image to be identified is determined based on the media asset information; Generate the recognition result of the image to be recognized; The result display area is set in the user interface according to the source type of the image to be identified; Extract the coordinate parameters of the result display area in the user interface, the coordinate parameters including spatial location and area shape data; If the source type of the image to be identified is a 2D source in the first type or a 3D source in the second type, extract the identification mark position from the identification result; Obtain the spatial location of the result display area, including the coordinates of the upper left corner and the upper right corner of the result display area; Based on the location of the identification mark and the spatial location, the coordinates of the upper left and upper right corners of the identification mark in the user interface are calculated so that the identification result is displayed in the result display area.

6. A virtual reality device, characterized in that, include: The monitor is configured to display the user interface; The controller is configured as follows: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the media asset information of the image to be identified is extracted, and the source type of the image to be identified is determined based on the media asset information; Generate the recognition result of the image to be recognized; The result display area is set in the user interface according to the source type of the image to be identified; Extract the coordinate parameters of the result display area in the user interface, the coordinate parameters including spatial location and area shape data; If the source type of the image to be identified is a 360-degree panoramic source in the first type, extract the identification marker position from the identification result; Convert the location of the identification marker into latitude and longitude; Obtain the region shape data of the result display area, wherein the region shape data includes the radius of the display sphere; Based on the latitude and longitude and the shape data of the region, the position coordinates of the identification mark in the user interface are calculated so that the identification result is displayed in the result display area.

7. A virtual reality device, characterized in that, include: The monitor is configured to display the user interface; The communicator is configured to connect to the server; The controller is configured as follows: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the source type of the image to be identified is detected; If the source image of the image to be identified is of the first type, the original source image is extracted as the image to be identified; If the source image of the image to be identified is of the second type, extract the half-side image corresponding to the left or right display in the source image as the image to be identified; The communicator sends an image recognition request to the server. Receive the recognition result fed back by the server based on the image to be recognized; The recognition result is displayed in the user interface according to the source type of the image to be recognized.

8. A VR scene image recognition method, characterized in that, Applied to virtual reality devices, the method includes: Obtain the control commands input by the user to initiate image recognition; In response to the control command, the source type of the image to be identified is detected; If the source image of the image to be identified is of the first type, the original source image is extracted as the image to be identified; Image recognition is performed on the original source image to generate recognition results; If the source image of the image to be identified is of the second type, extract the half-side image corresponding to the left or right display in the source image as the image to be identified; Image recognition is performed on a half-side image of the source image to generate a recognition result; The recognition results are displayed in the user interface according to the source type of the image to be recognized.

Citation Information

Patent Citations

  • Method and device for applying 3D image

    CN106971129A

  • Image recognition method, device and system

    CN111931835A