Method for realizing interaction in a metaverse online exhibition hall based on VR technology
By applying VR technology in the online exhibition hall of the Metaverse, users can efficiently query exhibit information and interact with virtual users, solving the problems of low query efficiency and poor interaction experience in the existing technology, and achieving a better user experience.
Patent Information
- Application Number
- CN202410943164.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2044-07-15
AI Technical Summary
In the Metaverse Online Exhibition Hall, users are inefficient when querying exhibit information and cannot interact with other users or virtual users, resulting in poor user experience.
Through a VR technology-based method, users can select exhibits in the virtual exhibition hall through a head-mounted display device, identify exhibit information using a pre-trained exhibit recognition model, and realize interaction with virtual users through action collection devices and voice recognition technology.
It improves the efficiency of querying exhibit information, enhances the user's interactive experience in the virtual space, and improves the user's overall experience.
Smart Images

Figure CN118915907B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computer technology, and more particularly to a method for realizing metaverse online exhibition hall interaction based on VR technology. Background Art
[0002] The metaverse is a virtual space that can interact with the real world online. The metaverse is a virtual enhanced physical reality, a 3D virtual space based on the future Internet that presents convergent and physically persistent characteristics and has connection perception and sharing characteristics. Common virtual reality devices include smartphones, smart glasses, virtual reality head-mounted displays, etc. Users can interact with other users through the metaverse exhibition hall. Currently, when users browse exhibits in the metaverse exhibition hall, the commonly adopted method is: obtaining an exhibit catalog and viewing the information of the exhibits to be queried one by one from the exhibit catalog.
[0003] However, it is found in practice that when browsing exhibits in the above manner, there is often the following technical problem 1: When viewing the information of the exhibits to be queried one by one, it takes a long time to determine the information of the exhibits to be queried, and the query efficiency is low.
[0004] In the process of adopting technical solutions to solve the above technical problem 1, there is often the following technical problem 2: Users cannot interact with other users or virtual users (e.g., NPCs) in the virtual space, and the user experience is poor. For these problems of the above technical problem 2, the conventional solutions are generally: interacting with other users or virtual users in text form. However, the above conventional solutions still have the following problems: When interacting in text form, relevant text input devices need to be used, resulting in poor practicality and convenience of the devices, and it takes a long time to input text when interacting in text form, and the interaction efficiency is low.
[0005] In the process of adopting technical solutions to solve the above technical problem 1, there is often the following technical problem 3: Users' attention to different exhibits is different, and the attention of users to virtual exhibits is not considered, resulting in poor user experience. For these problems of the above technical problem 3, the conventional solutions are generally: determining the preferences of users and determining the exhibited exhibit models through the preferences. However, the above conventional solutions still have the following problems: It is not considered that users' preferences will change over time, resulting in the determined exhibit models not conforming to users' preferences, and the user experience is poor.
[0006] The above information disclosed in this background art section is only used to enhance the understanding of the background of the inventive concept, and thus, it may include information that does not form the prior art known to those of ordinary skill in the art in this country. Summary of the Invention
[0007] This disclosure is in part for introducing concepts in a concise form, which will be described in detail in the following detailed implementation section. This disclosure is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] Some embodiments of this disclosure propose methods, devices, electronic devices, and computer-readable media for realizing metaverse online exhibition hall interaction based on VR technology to solve one or more of the technical problems mentioned in the above background section.
[0009] In a first aspect, some embodiments of this disclosure provide a method for realizing metaverse online exhibition hall interaction based on VR technology. The method includes: in response to detecting a startup operation on the above-mentioned head-mounted display device, obtaining a virtual exhibition hall space; determining a user starting coordinate corresponding to a target user; generating a user virtual model corresponding to the target user, and loading the user virtual model at the user starting coordinate; in response to the action collection device collecting a selection operation of the target user on any virtual exhibit model in the virtual exhibition hall space, intercepting an exhibit image of the virtual exhibit model; inputting the exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information; determining exhibit information according to the exhibit recognition information, and displaying the exhibit information on the head-mounted display device.
[0010] In a second aspect, some embodiments of this disclosure provide a device for realizing metaverse online exhibition hall interaction based on VR technology. The device includes: an obtaining unit configured to obtain a virtual exhibition hall space in response to detecting a startup operation on the above-mentioned head-mounted display device; a first determination unit configured to determine a user starting coordinate corresponding to a target user; a generating unit configured to generate a user virtual model corresponding to the target user, and load the user virtual model at the user starting coordinate; an intercepting unit configured to intercept an exhibit image of the virtual exhibit model in response to the action collection device collecting a selection operation of the target user on any virtual exhibit model in the virtual exhibition hall space; an input unit configured to input the exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information; a second determination unit configured to determine exhibit information according to the exhibit recognition information, and display the exhibit information on the head-mounted display device.
[0011] In a third aspect, some embodiments of this disclosure provide an electronic device, including: one or more processors; a storage device storing one or more programs thereon, and when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect above.
[0012] In a fourth aspect, some embodiments of the present disclosure provide a computer-readable medium having a computer program stored thereon, wherein when the program is executed by a processor, the method described in any implementation manner of the first aspect above is implemented.
[0013] The above various embodiments of the present disclosure have the following beneficial effects: Through the method for realizing metaverse online exhibition hall interaction based on VR technology in some embodiments of the present disclosure, the time for querying exhibit information is reduced, and the query efficiency is improved. Specifically, the reason for the need to spend a long time determining the exhibit information to be queried and the low query efficiency is that when viewing the exhibit information to be queried one by one, it takes a long time to determine the exhibit information to be queried, and the query efficiency is low. Based on this, the method for realizing metaverse online exhibition hall interaction based on VR technology in some embodiments of the present disclosure, first, in response to detecting a startup operation on the above-mentioned head-mounted display device, obtain a virtual exhibition hall space. Thus, a metaverse exhibition hall can be obtained. Second, determine the user starting coordinate corresponding to the target user. Thus, the generation coordinate of the user can be determined. Then, generate the user virtual model corresponding to the above-mentioned target user, and load the user virtual model at the above-mentioned user starting coordinate. Thus, the user model can be loaded. After that, in response to the above-mentioned action acquisition device collecting a selection operation of the above-mentioned target user on any virtual exhibit model in the above-mentioned virtual exhibition hall space, intercept the exhibit image of the above-mentioned virtual exhibit model. Thus, the image of the exhibit to be queried can be determined. Then, input the above-mentioned exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information. Thus, the exhibit information of the required exhibit can be determined. Finally, according to the above-mentioned exhibit recognition information, determine the exhibit information and display the above-mentioned exhibit information on the above-mentioned head-mounted display device. Thus, the recognition of the exhibit is completed, the time for querying exhibit information is reduced, and the query efficiency is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the various embodiments of the present disclosure will become more apparent. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the elements and elements are not necessarily drawn to scale.
[0015] Figure 1 is a flowchart of some embodiments of a method for realizing metaverse online exhibition hall interaction based on VR technology according to the present disclosure;
[0016] Figure 2 is a schematic structural diagram of some embodiments of a device for realizing metaverse online exhibition hall interaction based on VR technology according to the present disclosure;
[0017] Figure 3It is a schematic structural diagram of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed implementation manners
[0018] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0019] In addition, it should be noted that for the sake of convenience of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.
[0020] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence relationship of the functions performed by these devices, modules or units.
[0021] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0022] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0023] The present disclosure will be described in detail below with reference to the drawings and in combination with the embodiments.
[0024] Figure 1 Flow 100 of some embodiments of a method for implementing metaverse online exhibition hall interaction based on VR technology according to the present disclosure is shown. The method for implementing metaverse online exhibition hall interaction based on VR technology includes the following steps:
[0025] Step 101, in response to detecting a startup operation on the above-mentioned head-mounted display device, obtain a virtual exhibition hall space.
[0026] In some embodiments, the execution subject of the method for realizing the interaction of the metaverse online exhibition hall based on VR technology (such as a head-mounted display device) may, in response to detecting a startup operation acting on the above-mentioned head-mounted display device, obtain a virtual exhibition hall space. Among them, the above-mentioned method for realizing the interaction of the metaverse online exhibition hall based on VR technology is applied to the head-mounted display device. The above-mentioned head-mounted display device is connected to an action acquisition device. The above-mentioned head-mounted display device may be a VR glasses. The above-mentioned action acquisition device may be a device for collecting the behavioral actions of users. The above-mentioned head-mounted display device may include: a sound collection device and a camera. The above-mentioned sound collection device may be used to collect the speech of users. The above-mentioned camera may be used to collect the eye information of users (such as pupil images). The above-mentioned virtual exhibition hall space may be a pre-generated virtual space for displaying virtual exhibits.
[0027] In practice, the above-mentioned virtual exhibition hall space may be generated through the following steps:
[0028] The first step is to obtain an initial virtual space and a set of virtual exhibit models. Among them, the above-mentioned initial virtual space may be a pre-generated virtual space. The virtual exhibit models in the above-mentioned set of virtual exhibit models may be virtual models of exhibits preset for exhibition.
[0029] The second step is to perform the following processing steps for each virtual exhibit model in the above-mentioned set of virtual exhibit models:
[0030] The first processing step is to determine the exhibit positioning information corresponding to the above-mentioned virtual exhibit model. Among them, the above-mentioned exhibit positioning information may be the coordinate information of the virtual exhibit model preset in the above-mentioned initial virtual space.
[0031] The second processing step is to generate a virtual exhibition stand at the position corresponding to the above-mentioned exhibit positioning information in the above-mentioned initial virtual space. Among them, the above-mentioned virtual exhibition stand may be a virtual exhibition stand for displaying the virtual exhibit model.
[0032] The third processing step is to load the above-mentioned virtual exhibit model into the above-mentioned virtual exhibition stand;
[0033] The third step is to determine the initial virtual space loaded with each virtual exhibit model as the virtual exhibition hall space.
[0034] Considering the problems of the above-mentioned conventional solutions, in the face of the above-mentioned technical problem 3: the attention of users to different exhibits is different, and the attention of users to virtual exhibits is not considered, resulting in a poor user experience. Combining the current technical situation, the following solutions can be decided to be adopted:
[0035] Optionally, after step 101, the following steps are further included:
[0036] First step, control the camera included in the above-mentioned head-mounted display device to perform pupil recognition processing on the target user to generate a pupil recognition result.
[0037] In some embodiments, the above-mentioned execution entity may control the camera included in the above-mentioned head-mounted display device to perform pupil recognition processing on the target user to generate a pupil recognition result.
[0038] Second step, in response to the above-mentioned pupil recognition result satisfying the second preset condition, obtain the user historical behavior information set of the above-mentioned target user.
[0039] In some embodiments, the above-mentioned execution entity may, in response to the above-mentioned pupil recognition result satisfying the second preset condition, obtain the user historical behavior information set of the above-mentioned target user. Among them, the user historical behavior information in the above-mentioned user historical behavior information set includes the historical behavior time point. The above-mentioned second preset condition may indicate that there is historical access data for the above-mentioned target user.
[0040] Third step, for each user historical behavior information in the above-mentioned user historical behavior information set, perform user preference prediction processing on the above-mentioned target user according to the above-mentioned user historical behavior information to obtain a user preference label group.
[0041] In some embodiments, the above-mentioned execution entity may, for each user historical behavior information in the above-mentioned user historical behavior information set, perform user preference prediction processing on the above-mentioned target user according to the above-mentioned user historical behavior information to obtain a user preference label group. Among them, the user preference labels in the above-mentioned user preference label group represent the preferences of the target user for the virtual exhibit model. In practice, the above-mentioned execution entity may use a collaborative filtering algorithm to perform user preference prediction processing on the above-mentioned user historical behavior information.
[0042] Fourth step, perform a merging process on the obtained user preference label groups to generate a merged user preference label group.
[0043] In some embodiments, the above-mentioned execution entity may perform a merging process on the obtained user preference label groups to generate a merged user preference label group.
[0044] Fifth step, perform a deduplication process on each merged user preference label included in the above-mentioned merged user preference label group to generate a deduplicated user preference label group.
[0045] In some embodiments, the above-mentioned execution entity may perform a deduplication process on each merged user preference label included in the above-mentioned merged user preference label group to generate a deduplicated user preference label group.
[0046] Step 6: For each de-duplicated user preference tag in the above-mentioned de-duplicated user preference tag group, perform the following determination steps:
[0047] The first determination step: Determine the target time point and tag frequency corresponding to the above-mentioned de-duplicated user preference tag.
[0048] In some embodiments, the above-mentioned execution entity may determine the target time point and tag frequency corresponding to the above-mentioned de-duplicated user preference tag. Among them, the above-mentioned target time point may be the historical behavior time point included in the user historical behavior information corresponding to the de-duplicated user preference tag. The above-mentioned tag frequency may be the number of times the above-mentioned de-duplicated user preference tag appears in the above-mentioned combined user preference tag group.
[0049] The second determination step: Determine the preference tag weight corresponding to the above-mentioned de-duplicated user preference tag according to the above-mentioned target time point and tag frequency.
[0050] In some embodiments, the above-mentioned execution entity may determine the preference tag weight corresponding to the above-mentioned de-duplicated user preference tag according to the above-mentioned target time point and tag frequency.
[0051] Step 7: According to the determined preference tag weights of each, perform a sorting process on each de-duplicated user preference tag included in the above-mentioned de-duplicated user preference tag group to generate a de-duplicated user preference tag sequence.
[0052] In some embodiments, the above-mentioned execution entity may perform a sorting process on each de-duplicated user preference tag included in the above-mentioned de-duplicated user preference tag group according to the determined preference tag weights of each to generate a de-duplicated user preference tag sequence. Among them, the above-mentioned sorting process may be to sort each de-duplicated user preference tag in descending order according to the corresponding preference tag weight.
[0053] Step 8: According to the above-mentioned de-duplicated user preference tag sequence, perform a position swapping process on each virtual exhibit model in the above-mentioned virtual exhibition hall space to generate a swapped virtual exhibition hall space as the virtual exhibition hall space.
[0054] In some embodiments, the above-mentioned execution entity may perform a position swapping process on each virtual exhibit model in the above-mentioned virtual exhibition hall space according to the above-mentioned de-duplicated user preference tag sequence to generate a swapped virtual exhibition hall space as the virtual exhibition hall space.
[0055] The above first step - eighth step, as an inventive point of the embodiment of the present disclosure, solves the third technical problem mentioned in the background art, "Users have different attentions to different exhibits, and the attention of users to virtual exhibits is not considered, resulting in a poor user experience." The reasons for the poor user experience are as follows: Users have different attentions to different exhibits, and the attention of users to virtual exhibits is not considered, resulting in a poor user experience. If the above factors are solved, the effect of improving the user experience can be achieved. To achieve this effect, the present disclosure first controls the camera included in the above head - mounted display device to perform pupil recognition processing on the target user to generate a pupil recognition result. Thus, it can be determined whether the target user has entered the virtual exhibit space. Second, in response to the pupil recognition result satisfying the second preset condition, the user historical behavior information set of the above target user is obtained. Thus, the historical behavior of the target user can be obtained. Third, for each user historical behavior information in the above user historical behavior information set, according to the above user historical behavior information, user preference prediction processing is performed on the above target user to obtain a user preference tag group. Thus, the user preference of the target user can be determined. Fourth, the obtained user preference tag groups are merged to generate a merged user preference tag group; duplicate removal processing is performed on each merged user preference tag included in the above merged user preference tag group to generate a duplicate - removed user preference tag group. Thus, duplicate removal of each preference tag can be performed. Fifth, for each duplicate - removed user preference tag in the above duplicate - removed user preference tag group, the following determination steps are executed: Determine the target time point and tag frequency corresponding to the above duplicate - removed user preference tag; according to the above target time point and tag frequency, determine the preference tag weight corresponding to the above duplicate - removed user preference tag. Thus, the weight corresponding to each preference tag can be determined. Sixth, according to the determined preference tag weights, sorting processing is performed on each duplicate - removed user preference tag included in the above duplicate - removed user preference tag group to generate a duplicate - removed user preference tag sequence; according to the above duplicate - removed user preference tag sequence, position swapping processing is performed on each virtual exhibit model in the above virtual exhibition hall space to generate a swapped virtual exhibition hall space as the virtual exhibition hall space. Thus, the position of the exhibited virtual exhibit models can be swapped through the preference tags and the corresponding weights, so that users can view the virtual exhibit models with higher user attention, improving the user experience.
[0056] Step 102, determine the user starting coordinate corresponding to the target user.
[0057] In some embodiments, the above execution subject can determine the user starting coordinate corresponding to the target user. Among them, the above user starting coordinate can be the coordinate at which the target user is loaded into the above virtual exhibition hall space. The above target user can be the user wearing the above head - mounted display device.
[0058] In practice, the above-mentioned execution entity can determine the user starting coordinates corresponding to the target user through the following steps:
[0059] First step, obtain the virtual space configuration information corresponding to the initial virtual space. Among them, the above-mentioned virtual space configuration information may be the parameter information for generating the above-mentioned initial virtual space.
[0060] Second step, perform parsing processing on the above-mentioned virtual space configuration information to generate a parsing result.
[0061] Third step, according to the above-mentioned parsing result, determine the virtual space starting coordinates included in the above-mentioned initial virtual space. Among them, the above-mentioned virtual space starting coordinates may be the starting coordinates set in the above-mentioned initial virtual space.
[0062] Fourth step, map the user position coordinates of the above-mentioned target user to the above-mentioned initial virtual space to obtain mapped coordinates.
[0063] Fifth step, perform matching processing on the above-mentioned virtual space starting coordinates and the above-mentioned mapped coordinates to determine the above-mentioned virtual space coordinates as the user starting coordinates. Among them, the above-mentioned matching processing may be to overlap the above-mentioned virtual space starting coordinates and the above-mentioned mapped coordinates to generate user starting coordinates.
[0064] Step 103, generate a user virtual model corresponding to the target user, and load the user virtual model at the user starting coordinates.
[0065] In some embodiments, the above-mentioned execution entity can generate a user virtual model corresponding to the above-mentioned target user, and load the above-mentioned user virtual model at the above-mentioned user starting coordinates. Among them, the above-mentioned user virtual model can be used to represent the target user in the virtual exhibit space. In practice, the above-mentioned execution entity can place the above-mentioned user virtual model at the above-mentioned user starting coordinates.
[0066] Optionally, after step 103, the following steps are further included:
[0067] First step, control the above-mentioned action acquisition device to collect the user behavior information of the above-mentioned target user in real time.
[0068] In some embodiments, the above-mentioned execution entity can control the above-mentioned action acquisition device to collect the user behavior information of the above-mentioned target user in real time.
[0069] Second step, perform behavior recognition processing on the above-mentioned user behavior information to generate user behavior recognition information.
[0070] In some embodiments, the above-mentioned execution entity may perform behavior recognition processing on the above-mentioned user behavior information to generate user behavior recognition information. In practice, the above-mentioned user behavior information may be input into a pre-trained behavior recognition model to obtain user behavior recognition information.
[0071] The third step is to select a preset user behavior model corresponding to the above-mentioned user behavior recognition information from a preset user behavior model library.
[0072] In some embodiments, the above-mentioned execution entity may select a preset user behavior model corresponding to the above-mentioned user behavior recognition information from a preset user behavior model library.
[0073] The fourth step is to perform behavior adjustment processing on the above-mentioned user virtual model according to the above-mentioned preset user behavior model.
[0074] In some embodiments, the above-mentioned execution entity may perform behavior adjustment processing on the above-mentioned user virtual model according to the above-mentioned preset user behavior model.
[0075] Step 104: In response to the action acquisition device collecting a selection operation of the target user on any virtual exhibit model in the virtual exhibition hall space, intercept the exhibit image of the virtual exhibit model.
[0076] In some embodiments, the above-mentioned execution entity may, in response to the action acquisition device collecting the above-mentioned selection operation of the above-mentioned target user on any virtual exhibit model in the above-mentioned virtual exhibition hall space, intercept the exhibit image of the above-mentioned virtual exhibit model. Among them, the above-mentioned selection operation may include, but is not limited to, a click operation. The above-mentioned exhibit image may be an image for displaying the virtual exhibit model.
[0077] Step 105: Input the exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information.
[0078] In some embodiments, the above-mentioned execution entity may input the above-mentioned exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information. Among them, the above-mentioned exhibit recognition model may be a pre-trained classification model that takes the exhibit image as the input and the exhibit recognition information as the output. The above-mentioned exhibit recognition information may uniquely represent a certain virtual exhibit model. For example, the above-mentioned exhibit recognition information may be the code of the virtual exhibit model.
[0079] Optionally, the above-mentioned exhibit recognition model may be trained through the following steps:
[0080] The first step is to obtain a sample set.
[0081] In some embodiments, the above-mentioned execution entity may obtain a sample set. Among them, the samples in the above-mentioned sample set include sample exhibit images and sample exhibit identification information corresponding to the above-mentioned sample exhibit images.
[0082] The second step is to select a sample from the above-mentioned sample set.
[0083] In some embodiments, the above-mentioned execution entity may select a sample from the above-mentioned sample set. Here, the above-mentioned execution entity may randomly select a sample from the above-mentioned sample set.
[0084] The third step is to input the above-mentioned sample into the initial network model to obtain the exhibit identification information corresponding to the above-mentioned sample.
[0085] In some embodiments, the above-mentioned execution entity may input the above-mentioned sample into the initial network model to obtain the exhibit identification information corresponding to the above-mentioned sample. Among them, the above-mentioned initial neural network may be a classification model capable of obtaining exhibit identification information based on exhibit images. The above-mentioned initial neural network may be a classification model.
[0086] The fourth step is to determine the loss value between the above-mentioned exhibit identification information and the sample exhibit identification information included in the above-mentioned sample.
[0087] In some embodiments, the above-mentioned execution entity may determine the loss value between the above-mentioned exhibit identification information and the sample exhibit identification information included in the above-mentioned sample. In practice, the loss value between the above-mentioned exhibit identification information and the sample exhibit identification information included in the above-mentioned sample may be determined based on a preset loss function. For example, the above-mentioned preset loss function may be a cross-entropy loss function.
[0088] The fifth step is to adjust the network parameters of the above-mentioned initial network model in response to the above-mentioned loss value being greater than or equal to a preset threshold.
[0089] In some embodiments, the above-mentioned execution entity may adjust the network parameters of the above-mentioned initial network model in response to the above-mentioned loss value being greater than or equal to a preset threshold. Here, there is no limitation on the setting of the preset threshold. For example, the difference between the loss value and the preset threshold may be calculated to obtain a loss difference. On this basis, methods such as backpropagation and stochastic gradient descent may be used to forward the error value from the last layer of the model to adjust the parameters of each layer. Of course, according to needs, the method of network freezing (dropout) may also be adopted to keep the network parameters of some layers unchanged without adjustment, and no limitation is made on this.
[0090] Optionally, in response to the above-mentioned loss value being less than the above-mentioned preset threshold, the above-mentioned initial network model is determined as an exhibit identification model.
[0091] In some embodiments, the above-mentioned execution entity may determine the initial network model as the exhibit recognition model in response to the above-mentioned loss value being less than the above-mentioned preset threshold.
[0092] Step 106: Determine exhibit information according to the exhibit recognition information, and display the exhibit information on the head-mounted display device.
[0093] In some embodiments, the above-mentioned execution entity may determine exhibit information according to the above-mentioned exhibit recognition information, and display the above-mentioned exhibit information on the above-mentioned head-mounted display device. In practice, the exhibit information corresponding to the above-mentioned exhibit recognition information may be obtained from the target database. Wherein, the above-mentioned target database may be a database storing exhibit information. The above-mentioned exhibit information may be information for describing a virtual exhibit model.
[0094] Considering the problems of the above-mentioned conventional solutions, in the face of the above-mentioned technical problem 2: The user cannot interact with other users or virtual users in the virtual space. Combining the existing technical status, the following solution can be decided to be adopted:
[0095] Optionally, after step 106, the following steps are further included:
[0096] First step: Real-time detect whether the user virtual model corresponding to the above-mentioned target user meets the first preset condition.
[0097] In some embodiments, the above-mentioned execution entity may real-time detect whether the user virtual model corresponding to the above-mentioned target user meets the first preset condition. Wherein, the above-mentioned first preset condition may be that the above-mentioned user virtual model is within a preset range centered on the location of the virtual user.
[0098] Second step: In response to detecting that the above-mentioned user virtual model meets the above-mentioned first preset condition, control the sound collection device included in the above-mentioned head-mounted display device to real-time collect the audio data of the above-mentioned target user.
[0099] In some embodiments, the above-mentioned execution entity may, in response to detecting that the above-mentioned user virtual model meets the above-mentioned first preset condition, control the sound collection device included in the above-mentioned head-mounted display device to real-time collect the audio data of the above-mentioned target user.
[0100] Third step: Preprocess the above-mentioned audio data to generate preprocessed audio data. Wherein, the above-mentioned preprocessing may be noise reduction processing on the above-mentioned audio data.
[0101] Fourth step: Perform audio recognition processing on the above-mentioned preprocessed audio data to generate audio recognition information.
[0102] In some embodiments, the above-mentioned execution entity may perform audio recognition processing on the preprocessed audio data to generate audio recognition information. Among them, the above-mentioned audio recognition processing may be to input the preprocessed audio data into a pre-trained audio recognition model to obtain audio recognition information.
[0103] Step 5: Input the above-mentioned audio recognition information into a pre-trained large language model to obtain an audio reply message.
[0104] In some embodiments, the above-mentioned execution entity may input the above-mentioned audio recognition information into a pre-trained large language model to obtain an audio reply message. Among them, the above-mentioned large language model may be a large language model deployed in a local server.
[0105] Step 6: Input the above-mentioned audio reply message into a pre-trained speech generation model to obtain a reply audio.
[0106] In some embodiments, the above-mentioned execution entity may input the above-mentioned audio reply message into a pre-trained speech generation model to obtain a reply audio.
[0107] Step 7: Perform playback processing on the above-mentioned reply audio and control a preset user model to match the above-mentioned reply audio.
[0108] In some embodiments, the above-mentioned execution entity may perform playback processing on the above-mentioned reply audio and control a preset user model to match the above-mentioned reply audio. In practice, the facial model of the preset user model may be matched according to the reply audio.
[0109] The above first step to seventh step, as an inventive point of the embodiment of the present disclosure, solves the third technical problem mentioned in the background art, "Users cannot interact with other users or virtual users in the virtual space, and the user experience is poor." The reasons for the poor user experience are as follows: Users cannot interact with other users or virtual users in the virtual space, and the user experience is poor. If the above factors are solved, the effect of improving the user experience can be achieved. To achieve this effect, the present disclosure first, in real time, detects whether the user virtual model corresponding to the target user meets the first preset condition. Thus, it can be determined whether the user is within the conversation range of the virtual user. Second, in response to detecting that the user virtual model meets the first preset condition, controls the sound collection device included in the head-mounted display device to collect the audio data of the target user in real time. Thus, the conversation audio of the user can be recorded. Third, preprocesses the audio data to generate preprocessed audio data. Thus, it is convenient to recognize the conversation audio of the user. Fourth, performs audio recognition processing on the preprocessed audio data to generate audio recognition information; inputs the audio recognition information into a pre-trained large language model to obtain audio reply information. Thus, the reply content of the virtual user can be determined. Fifth, inputs the audio reply information into a pre-trained speech generation model to obtain a reply audio. Thus, a reply audio can be generated. Sixth, performs playback processing on the reply audio and controls a preset user model to match the reply audio. Thus, the conversation between the target user and the virtual user is completed, and the user experience is improved.
[0110] The above-mentioned various embodiments of the present disclosure have the following beneficial effects: Through the method for realizing metaverse online exhibition hall interaction based on VR technology in some embodiments of the present disclosure, the time for querying exhibit information is reduced, and the query efficiency is improved. Specifically, the reason for the need to spend a long time to determine the exhibit information to be queried and the low query efficiency is that when viewing the exhibit information to be queried one by one, it takes a long time to determine the exhibit information to be queried, and the query efficiency is low. Based on this, in some embodiments of the method for realizing metaverse online exhibition hall interaction based on VR technology of the present disclosure, first, in response to detecting a startup operation on the above-mentioned head-mounted display device, a virtual exhibition hall space is obtained. Thus, a metaverse exhibition hall can be obtained. Secondly, determine the user starting coordinate corresponding to the target user. Thus, the generation coordinate of the user can be determined. Then, generate the user virtual model corresponding to the above-mentioned target user, and load the user virtual model at the above-mentioned user starting coordinate. Thus, the user model can be loaded. After that, in response to the above-mentioned motion capture device capturing a selection operation of the above-mentioned target user on any virtual exhibit model in the above-mentioned virtual exhibition hall space, capture the exhibit image of the above-mentioned virtual exhibit model. Thus, the image of the exhibit to be queried can be determined. Then, input the above-mentioned exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information. Thus, the exhibit information of the required exhibit can be determined. Finally, according to the above-mentioned exhibit recognition information, determine the exhibit information and display the above-mentioned exhibit information on the above-mentioned head-mounted display device. Thus, the recognition of the exhibit is completed, the time for querying exhibit information is reduced, and the query efficiency is improved.
[0111] Further referring to Figure 2 , as an implementation of the methods shown in the above figures, the present disclosure provides some embodiments of a device for realizing metaverse online exhibition hall interaction based on VR technology. These device embodiments correspond to Figure 1 the method embodiments shown, and the device for realizing metaverse online exhibition hall interaction based on VR technology can be specifically applied to various electronic devices.
[0112] Such as Figure 2As shown, the device 200 for realizing the interaction of the metaverse online exhibition hall based on VR technology in some embodiments includes: an acquisition unit 201, a first determination unit 202, a generation unit 203, a capture unit 204, an input unit 205, and a second determination unit 206. Among them, the acquisition unit 201 is configured to acquire a virtual exhibition hall space in response to detecting a startup operation acting on the above-mentioned head-mounted display device; the first determination unit 202 is configured to determine the user starting coordinates corresponding to the target user; the generation unit 203 is configured to generate a user virtual model corresponding to the above-mentioned target user, and load the user virtual model at the above-mentioned user starting coordinates; the capture unit 204 is configured to capture an exhibit image of the above-mentioned virtual exhibit model in response to the above-mentioned motion capture device capturing a selection operation of the above-mentioned target user acting on any virtual exhibit model in the above-mentioned virtual exhibition hall space; the input unit 205 is configured to input the above-mentioned exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information; the second determination unit 206 is configured to determine exhibit information according to the above-mentioned exhibit recognition information, and display the above-mentioned exhibit information on the above-mentioned head-mounted display device.
[0113] It can be understood that the units described in the device 200 for realizing the interaction of the metaverse online exhibition hall based on VR technology correspond to the respective steps in the method described in the reference Figure 1 Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 for realizing the interaction of the metaverse online exhibition hall based on VR technology and the units included therein, and will not be repeated here.
[0114] Next, refer to Figure 3 , which shows a schematic structural diagram of an electronic device 300 suitable for implementing some embodiments of the present disclosure. The electronic devices in some embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 3 The electronic device shown is only an example and should not impose any limitations on the functions and usage ranges of the embodiments of the present disclosure.
[0115] As Figure 3As shown, the electronic device 300 may include a processing device 301 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in the read-only memory (ROM) 302 or a program loaded from the storage device 308 into the random access memory (RAM) 303. In the RAM 303, various programs and data required for the operation of the electronic device 300 are also stored. The processing device 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. The input / output (I / O) interface 305 is also connected to the bus 304.
[0116] Generally, the following devices may be connected to the I / O interface 305: an input device 306 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 307 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 308 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 309. The communication device 309 may allow the electronic device 300 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 3 an electronic device 300 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. More or fewer devices may be implemented or had alternatively. Figure 3 Each block shown in may represent one device or, as needed, multiple devices.
[0117] Specifically, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from the network through the communication device 309, or installed from the storage device 308, or installed from the ROM 302. When the computer program is executed by the processing device 301, the above functions defined in the methods of some embodiments of the present disclosure are executed.
[0118] It should be noted that the computer-readable media described in some embodiments of the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take many forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0119] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0120] The above computer-readable medium may be included in the above electronic device; or may exist separately without being assembled into the electronic device. The above computer-readable medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device is caused to: in response to detecting a startup operation applied to the above head-mounted display device, obtain a virtual exhibition hall space. Determine a user starting coordinate corresponding to the target user. Generate a user virtual model corresponding to the above target user, and load the user virtual model at the above user starting coordinate. In response to the above action collection device collecting a selection operation applied by the above target user to any virtual exhibit model in the above virtual exhibition hall space, capture an exhibit image of the above virtual exhibit model. Input the above exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information. Determine exhibit information according to the above exhibit recognition information, and display the above exhibit information on the above head-mounted display device.
[0121] Computer program code for performing the operations of some embodiments of the present disclosure may be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., by connecting through the Internet using an Internet service provider).
[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0123] The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit, a first determination unit, a generation unit, a capture unit, an input unit, and a second determination unit. Among them, the names of these units do not constitute a limitation to the unit itself in some cases. For example, the acquisition unit can also be described as "a unit that acquires the virtual exhibition hall space in response to detecting a startup operation acting on the above-mentioned head-mounted display device".
[0124] The functions described above can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0125] The above description is only some preferred embodiments of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.
Claims
1. A method for realizing the interaction of Metaverse online exhibition hall based on VR technology, applied to a head-mounted display device, wherein: The head mounted display device is connected to a motion acquisition device, and the method comprises: In response to detecting a start-up operation acting on the head mounted display device, acquiring a virtual exhibition hall space; Determine the user starting coordinates corresponding to the target user; Generate a user virtual model corresponding to the target user, and load the user virtual model at the user starting coordinates; In response to the action acquisition device acquiring a selection operation of the target user on any virtual exhibit model in the virtual exhibition hall space, capturing an exhibit image of the virtual exhibit model; Inputting the exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information; The exhibit recognition model is trained through the following steps: Acquire a sample set, wherein the samples in the sample set include sample exhibit images and sample exhibit identification information corresponding to the sample exhibit images; Selecting a sample from the sample set; Inputting the sample into the initial network model to obtain exhibit identification information corresponding to the sample; determining a loss value between the exhibit identification information corresponding to the sample and the sample exhibit identification information included in the sample; In response to the loss value being greater than or equal to a preset threshold, adjusting a network parameter of the initial network model; In response to the loss value being less than the preset threshold, determining the initial network model as an exhibit recognition model; The exhibit information is determined according to the exhibit identification information, and the exhibit information is displayed on the head mounted display device.
2. The method according to claim 1, wherein: The virtual exhibition hall space is generated by the following steps: Obtaining an initial set of virtual space and virtual exhibit models; For each virtual exhibit model in the virtual exhibit model set, the following processing steps are performed: Determining exhibit location information corresponding to the virtual exhibit model; Generating a virtual exhibition stand at a position corresponding to the exhibit location information in the initial virtual space; Loading the virtual exhibit model into the virtual exhibition stand; The initial virtual space loaded with each virtual exhibit model is determined as the virtual exhibition hall space.
3. The method according to claim 1, wherein: The determining of the user starting coordinates corresponding to the target user includes: Acquire virtual space configuration information corresponding to the initial virtual space; Parsing the virtual space configuration information to generate a parsing result; Determining the virtual space starting coordinates included in the initial virtual space according to the analysis result; Mapping the user position coordinates of the target user into the initial virtual space to obtain mapping coordinates; The virtual space starting coordinates and the mapping coordinates are matched to determine the virtual space coordinates as the user starting coordinates.
4. The method according to claim 1, wherein: After generating the user virtual model corresponding to the target user and loading the user virtual model at the user starting coordinates, the method further includes: Controlling the action acquisition device to acquire user behavior information of the target user in real time; Performing behavior recognition processing on the user behavior information to generate user behavior recognition information; Selecting a preset user behavior model corresponding to the user behavior identification information from a preset user behavior model library; According to the preset user behavior model, behavior adjustment processing is performed on the user virtual model.
5. A device for realizing the interaction of Metaverse online exhibition hall based on VR technology, comprising: an acquisition unit configured to acquire a virtual exhibition hall space in response to detecting a start-up operation acting on the head mounted display device; A first determining unit is configured to determine a user starting coordinate corresponding to a target user; A generating unit, configured to generate a user virtual model corresponding to the target user, and load the user virtual model at the user starting coordinates; a capture unit configured to capture an exhibit image of the virtual exhibit model in response to the action capture device capturing a selection operation of the target user on any virtual exhibit model in the virtual exhibition hall space; The input unit is configured to input the exhibit image into a pre-trained exhibit recognition model to obtain exhibit recognition information; the input unit is further configured as follows: wherein the exhibit recognition model is trained by the following steps: Acquire a sample set, wherein the samples in the sample set include sample exhibit images and sample exhibit identification information corresponding to the sample exhibit images; Selecting a sample from the sample set; Inputting the sample into the initial network model to obtain exhibit identification information corresponding to the sample; determining a loss value between the exhibit identification information corresponding to the sample and the sample exhibit identification information included in the sample; In response to the loss value being greater than or equal to a preset threshold, adjusting a network parameter of the initial network model; In response to the loss value being less than the preset threshold, determining the initial network model as an exhibit recognition model; The second determining unit is configured to determine the exhibit information according to the exhibit identification information, and display the exhibit information in the head mounted display device.
6. An electronic device comprising: one or more processors; a storage device having one or more programs stored thereon; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 4.
7. A computer readable medium having a computer program stored thereon, wherein: When the program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Exhibition information watching system and method in virtual reality environment
CN114511229A
Cloud exhibition hall system based on element universe
CN117315132A
Information processing method, information processing system, and program
JP2024088938A