Method and apparatus for performing interaction on basis of virtual object, and electronic device and medium
By acquiring the positional information of the user's operating parts, and using smart devices to display virtual objects to obtain demand information and generate response information, the problem of users having difficulty making decisions in shopping categories with high cognitive barriers is solved, and an efficient and convenient shopping experience is achieved.
Patent Information
- Application Number
- PCT/CN2025/103975
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-30
- Filing Date
- 2025-06-26
- Publication Date
- 2026-02-05
AI Technical Summary
Users often struggle to quickly and accurately find the knowledge or products they want in unfamiliar areas, especially in online shopping scenarios such as home improvement, where high cognitive barriers require expert explanations or self-study to make purchasing decisions.
By acquiring the pose information of the user's operating parts, the system uses smart devices to display virtual objects in the virtual world, obtains user demand information based on the interaction detection results, generates and displays response information that matches the demand information, and provides shopping decision support by having virtual objects accompany the user throughout the entire process in the virtual world.
It improves the efficiency and experience of users shopping in categories with high cognitive barriers, provides reliable and personalized shopping decision-making support, and avoids the decline in user experience caused by the limitations of terminal device screens.
Smart Images

Figure CN2025103975_05022026_PF_FP_ABST
Abstract
Description
Interaction methods, devices, electronic devices, and media based on virtual objects
[0001] This application claims priority to Chinese Patent Application No. 202411035061.1, filed on July 30, 2024, the contents of which are incorporated herein by reference. Technical Field
[0002] This disclosure relates to the technical fields of artificial intelligence, smart devices, and intelligent customer service, and more specifically, to an interaction method, device, electronic device, and medium based on virtual objects. Background Technology
[0003] With the continuous development of computer technology, users can use online search engines to learn knowledge and online shopping platforms to shop online. In such search scenarios, if users know what they want, they can find the information they are looking for through browsing content and searching.
[0004] However, in areas unfamiliar to users, it is difficult for them to quickly and accurately find the knowledge they want, the functions they need, or the products they want to buy. For example, in online shopping scenarios, for categories that users are not familiar with, such as home improvement or medicine, it is difficult for them to quickly and clearly understand the products they intend to buy and find products that meet their shopping needs. Summary of the Invention
[0005] In view of this, the present disclosure provides an interaction method, device, electronic device and medium based on virtual objects.
[0006] One aspect of this disclosure provides an interaction method based on virtual objects, comprising: acquiring pose information of a user's operating part in the real world, wherein the pose information is determined based on an image captured by a smart device, the smart device also being used to display a virtual world to the user, the virtual world including a virtual screen and virtual objects belonging to the same application software as the virtual screen; determining an interaction detection result for the virtual object based on the pose information; in response to determining that the interaction detection result indicates an operation on the virtual object is triggered, acquiring user demand information, wherein the user demand information includes demand information for the virtual screen; generating response information matching the user demand information based on the user demand information; and displaying the response information.
[0007] According to embodiments of this disclosure, the pose information includes first pose information; determining the interaction detection result for the virtual object based on the pose information includes: if the first pose information matches the preset pose information, determining the interaction detection result as triggering an operation on the virtual object; if the first pose information does not match the preset pose information, determining the interaction detection result as not triggering an operation on the virtual object.
[0008] According to embodiments of this disclosure, the pose information further includes first position information. Based on the pose information, determining the interaction detection result for the virtual object includes: determining a position detection result based on the first position information and the second position information of the virtual object, wherein the position detection result represents whether the operating part in the real world is in contact with the virtual object in the virtual world, and the second position information represents the position information of the virtual object in the real world; if the position detection result represents that the operating part is in contact with the virtual object and the first pose information matches the preset pose information, determining that the interaction detection result is a trigger for an operation on the virtual object.
[0009] According to embodiments of this disclosure, the method further includes: determining second position information of a virtual object based on third position information of a virtual image; or, acquiring pose information of a smart device, wherein the pose information includes fourth position information and second posture information; determining the second position information based on the fourth position information, the second posture information, and the user's arm span distance; the user's arm span distance characterizes the farthest distance when the user extends their arm.
[0010] According to embodiments of this disclosure, generating response information related to user demand information based on user demand information includes: determining user intent based on user demand information; and generating response information matching user demand information based on user intent.
[0011] According to embodiments of this disclosure, determining user intent based on user demand information includes: obtaining historical demand information and the current screen content of a virtual screen, wherein the current screen content includes at least one screen content corresponding to each of multiple shopping functions in the shopping platform; and using a large language model to determine user intent based on user demand information, historical demand information, and the current screen content.
[0012] According to embodiments of this disclosure, generating response information that matches user needs based on user intent includes: determining a response strategy that matches user intent; and generating response information based on user needs and the response strategy using a large language model.
[0013] According to embodiments of this disclosure, displaying response information includes: displaying the response information in text form to a user through a display window in a virtual world, wherein the position information of the display window is determined based on a second position information of the virtual object; or, converting the response information into speech data and generating a lip-shape change sequence matching the speech data; generating a virtual object speech response animation based on the lip-shape change sequence and model parameters of the virtual object; and outputting the speech data through a smart device and synchronously displaying the virtual object speech response animation in the virtual world.
[0014] According to embodiments of this disclosure, the method further includes: determining the user's emotional information based on user demand information and / or historical demand information; determining a virtual object display animation that matches the emotional information, wherein the virtual object display animation indicates changes in the virtual object's display actions and / or changes in its display expressions; and simultaneously displaying reply information and the virtual object display animation.
[0015] Another aspect of this disclosure provides an interactive device based on virtual objects, comprising: a first acquisition module for acquiring pose information of a user's operating part in the real world, wherein the pose information is determined based on an image captured by a smart device, and the smart device is also used to display a virtual world to the user, the virtual world including a virtual screen and virtual objects belonging to the same application software as the virtual screen; an interaction detection module for determining an interaction detection result for the virtual object based on the pose information; a second acquisition module for acquiring user demand information in response to determining that the interaction detection result indicates an operation on the virtual object, wherein the user demand information includes demand information for the virtual screen; a generation module for generating response information matching the user demand information based on the user demand information; and a display module for displaying the response information.
[0016] Another aspect of this disclosure provides an electronic device comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the methods described above.
[0017] Another aspect of this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed, are used to implement the methods described above.
[0018] Another aspect of this disclosure provides a computer program product including computer-executable instructions that, when executed, are used to implement the methods described above. Attached Figure Description
[0019] The above and other objects, features and advantages of this disclosure will become clearer from the following description of embodiments with reference to the accompanying drawings, in which:
[0020] Figure 1 schematically illustrates an application scenario where the interaction method and apparatus based on virtual objects disclosed herein can be applied;
[0021] Figure 2 schematically illustrates a flowchart of a virtual object-based interaction method according to an embodiment of the present disclosure;
[0022] Figure 3 schematically illustrates a scenario of a method for determining interaction detection results for the virtual object according to an embodiment of the present disclosure;
[0023] Figure 4 schematically illustrates a flowchart of a method for determining response information according to an embodiment of the present disclosure;
[0024] Figure 5 schematically illustrates a flowchart of displaying virtual object animation and reply information according to another embodiment of the present disclosure;
[0025] Figure 6 schematically illustrates a block diagram of a virtual object-based interactive device according to an embodiment of the present disclosure; and
[0026] Figure 7 schematically illustrates a block diagram of an electronic device suitable for implementing a virtual object-based interaction method and apparatus according to embodiments of the present disclosure. Detailed Implementation
[0027] The embodiments of the present disclosure will now be described with reference to the accompanying drawings. However, it should be understood that these descriptions are exemplary only and are not intended to limit the scope of the disclosure. In the following detailed description, numerous specific details are set forth to provide a thorough understanding of the embodiments of the present disclosure for ease of explanation. However, it will be apparent that one or more embodiments may be practiced without these specific details. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts of the present disclosure.
[0028] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this disclosure. The terms “comprising,” “including,” etc., as used herein indicate the presence of the stated features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.
[0029] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein are to be interpreted in a manner consistent with the context of this specification, and not in an idealized or overly rigid way.
[0030] When using expressions such as "at least one of A, B and C", they should generally be interpreted in accordance with the meaning that is commonly understood by those skilled in the art (e.g., "a system having at least one of A, B and C" should include, but is not limited to, a system having A alone, a system having B alone, a system having C alone, a system having A and B, a system having A and C, a system having B and C, and / or a system having A, B and C, etc.).
[0031] Related online shopping models can satisfy users' needs for quickly searching and purchasing products in categories with high standards and low cognitive barriers, such as fast-moving consumer goods, clothing, and information appliances (3C, a combination of computers, communications, and consumer electronics). For products with high standards and low cognitive barriers, users know exactly what they want and what they should want. However, for non-standard and high cognitive barrier categories, such as home improvement, users need a process of decoding their needs before making a purchase. That is, users need expert guidance or self-study (usually up to 3 months) to gradually understand what they should want and thus make a reliable purchasing decision.
[0032] Therefore, in order to provide users with reliable and personalized shopping decision-making support or recommendations throughout the entire shopping process, embodiments of this disclosure create a virtual object in the virtual world that belongs to the same application software as the virtual object. This virtual object can accompany the user throughout the entire shopping scenario and provide users with shopping decision-making support or recommendations when the user triggers an operation on the virtual object.
[0033] Specifically, embodiments of this disclosure provide an interaction method based on virtual objects, comprising: acquiring pose information of a user's operating part in the real world, wherein the pose information is determined based on an image captured by a smart device, and the smart device is also used to display a virtual world to the user, the virtual world including a virtual screen and virtual objects belonging to the same application software as the virtual screen; determining an interaction detection result for the virtual object based on the pose information; in response to determining that the interaction detection result indicates an operation on the virtual object is triggered, acquiring user demand information, the user demand information including demand information for the virtual screen; generating response information related to the user demand information based on the user demand information; and displaying the response information.
[0034] Figure 1 schematically illustrates an application scenario where the interaction method and apparatus based on virtual objects disclosed herein can be applied. It should be noted that Figure 1 is merely an example of an application scenario where embodiments of this disclosure can be applied, to help those skilled in the art understand the technical content of this disclosure, but does not imply that the embodiments of this disclosure cannot be used in other devices, systems, environments, or scenarios.
[0035] As shown in Figure 1, in embodiment 100 of the interaction method based on virtual objects, user 101 can wear a smart device 102 to present multiple virtual scenes 104 in the virtual world to user 101. It can be understood that when user 101 wears the smart device 102, he / she can view the virtual world through the smart device 102, and when user 101 does not wear the smart device 102, he / she can view the real world through his / her eyes.
[0036] Each virtual screen 104 can be a screen displayed by a smart device 102 after overlaying images, videos, application software pages, etc. onto the real world using virtual technology. For example, virtual technology includes, but is not limited to, the following implementation methods: Virtual Reality (VR), Augmented Reality (AR), Mixed Reality (MR), and Extended Reality (XR).
[0037] For example, virtual screen 104 can be a virtual page that is projected onto the homepage, search page, product browsing page, product details page, payment page, etc. of the shopping platform by smart device 102 and then displayed to user 101.
[0038] This embodiment 100 also includes a virtual object 103. The virtual object 103 may be a virtual object belonging to the same application software as the virtual screen, for example, it may be a virtual person or a virtual animal; or, when the virtual screen is a shopping platform screen, the virtual object may be the logo object of the shopping platform.
[0039] After wearing the smart device 102, user 101 can view virtual images 104 and virtual objects 103 on objects or backgrounds displayed in the real world. User 101 can interact with and control virtual images 104 and virtual objects 103 through gestures and other operations.
[0040] It should be noted that the virtual object-based interaction method provided in this disclosure can generally be implemented by a server. For example, the server can be a backend server that supports the application software to which the virtual object and virtual screen belong, such as the backend server of a shopping platform. Accordingly, the virtual object-based interaction device provided in this disclosure can generally be set up in a server.
[0041] Figure 2 schematically illustrates a flowchart of a virtual object-based interaction method according to an embodiment of the present disclosure.
[0042] As shown in Figure 2, the method includes operations S210~S250.
[0043] In operation S210, the user's operating part's position and pose information in the real world is acquired. The position and pose information is determined based on images captured by the smart device. The smart device is also used to display a virtual world to the user. The virtual world includes virtual images and virtual objects belonging to the same application software as the virtual images.
[0044] In operation S220, based on pose information, the interaction detection results for the virtual object are determined.
[0045] In operation S230, in response to determining the interaction detection result characterization triggering the operation on the virtual object, user demand information is obtained, wherein the user demand information includes demand information for the virtual screen.
[0046] In operation S240, a response message matching the user's needs information is generated based on the user's needs information.
[0047] When operating S250, display the reply information.
[0048] According to embodiments of this disclosure, the real world and the virtual world can be understood as the world that a user actually sees, and the world that is presented to the user through the virtual technology of smart devices. Users can view the virtual world, as well as virtual images and objects existing within it, through worn smart devices. Furthermore, although the user and their operating parts exist in the real world, smart devices can also display the user's operating parts in the virtual world.
[0049] For example, User A is wearing smart devices, such as VR glasses, while User B is next to User A without wearing any smart devices. User A can view virtual images, virtual objects, and User B through the VR glasses. User B can only see User A and cannot view the virtual images and virtual objects.
[0050] In one illustrative embodiment, after a user wears a smart device, the virtual object can be seen anywhere within the user's field of vision. As the user performs actions such as turning their head or looking down, which change the user's field of vision, the virtual object can also change accordingly.
[0051] According to embodiments of this disclosure, smart devices include, but are not limited to, VR devices, AR devices, MR devices, and XR devices. A smart device includes a camera unit for capturing images from at least one perspective; for example, the camera unit can capture images of the user's real-world environment, the user's operating body parts, etc.
[0052] According to embodiments of this disclosure, users can interact with virtual images and objects in the virtual world through the pose information of the operating part. The pose information of the operating part can indicate the position and posture of the operating part in the real world. For example, when the operating part is a hand, the pose information can include the position and posture of the hand, such as gestures.
[0053] In one specific embodiment, the smart device is equipped with a target detection algorithm that can perform target detection on the captured images to identify the posture information of the operating part. For example, the target detection algorithm can be a gesture recognition algorithm. The smart device can also analyze the captured images based on the position of the smart device and the position of the camera unit to determine the relative position of the operating part and the smart device / camera unit; then, combined with the position of the smart device, the position information of the operating part is determined.
[0054] In the embodiments of this disclosure, regarding operation S210, the server of the application software for the virtual screen and virtual object can obtain the pose information of the user's operating parts through interactive operation with the smart device. For example, the server can directly obtain the pose information from the smart device, or the server can also interact with the backend server of the smart device and obtain the pose information from the backend server.
[0055] According to embodiments of this disclosure, the virtual object displayed to the user belongs to the same application software as the virtual screen. After the user opens the application software, the virtual screen and virtual object are rendered and displayed to the user via a smart device. It can be understood that the virtual screen projected onto any page of the application software via the smart device is associated with the virtual object. When the application software is not open, neither the application software's virtual screen nor the virtual object is displayed.
[0056] For example, the virtual screen could be any shopping page within the shopping platform projected onto a real-world scene, and the virtual object could be a virtual shopping guide avatar of the platform. Once a user opens the shopping platform, they can continuously see the virtual object as they switch between different shopping pages projected onto the virtual screen. When the user exits the shopping platform, the virtual screen displayed on the smart device becomes unrelated to the platform, and consequently, the user can no longer see the virtual object.
[0057] According to embodiments of this disclosure, in operation S220, the interaction detection result for the virtual object can be determined based on pose information. Since the virtual object exists only in the virtual world, and the user and their operating parts cannot physically touch it, the user can perform various operations on the virtual object using the pose information of their operating parts to achieve interaction with the virtual object. In this case, the virtual object can also perform corresponding functions based on the user's operations. For example, it can perform the function of obtaining user request information.
[0058] The interaction detection results for virtual objects include whether an operation on the virtual object was triggered. In some embodiments, the interaction detection results may also include the specific operation performed by the user on the virtual object. For example, user operations on the virtual object include selecting the virtual object, touching the virtual object, clicking the virtual object, etc.
[0059] According to embodiments of this disclosure, in operation S230, when it is determined that a user has triggered an operation on a virtual object, user demand information is obtained. The user demand information includes demand information related to the virtual screen, such as demand information related to the content displayed on the virtual screen. When the virtual screen displays home decoration products, the user demand information could be "What does the AA symbol for home decoration products mean?" or "How to distinguish the quality of different seat materials?"
[0060] Since virtual objects and virtual screens belong to the same application software, as long as the user has user needs regarding the virtual screen, the virtual object can display response information related to the user's needs.
[0061] For operation S240, Natural Language Processing (NLP) can be used to generate response information that matches the user's needs. For example, NLP can be used to segment and extract features from the user's needs, and then relevant response information can be matched based on the extracted features.
[0062] According to embodiments of this disclosure, the response information can be various types of information such as product recommendations, function indexes, scenario introductions, and chat replies.
[0063] Regarding operation S250, since the user request information is obtained in response to an operation on a virtual object, the generated response information can be displayed through the virtual object to improve the user experience. For example, the response information can be displayed in an area near the virtual object.
[0064] Through embodiments of this disclosure, a smart device is used to display virtual images and virtual objects belonging to the same application software as the virtual images to the user, allowing the user to view the virtual images and virtual objects in an immersive way. When displaying virtual images and virtual objects to the user through a smart device, the user's gesture information triggers operations on the virtual objects, thereby obtaining user demand information. Subsequently, by generating and displaying response information matching the user demand information, the user's needs can be met at any time, such as providing the user with reliable and personalized shopping decision-making support or shopping decisions. Furthermore, since the virtual objects generate response information matching the user demand information in real time in the virtual world, the embodiments of this disclosure not only provide users with near-realistic auxiliary services but also avoid virtual objects obscuring the virtual image due to the screen limitations of mobile phones or terminal devices, thereby improving the user experience.
[0065] According to embodiments of this disclosure, the pose information of the operating part includes first pose information characterizing the pose of the operating part.
[0066] According to an embodiment of this disclosure, for operation S220, determining the interaction detection result for the virtual object based on pose information includes: if it is determined that the first pose information matches the preset pose information, determining that the interaction detection result is a triggered operation on the virtual object; if it is determined that the first pose information does not match the preset pose information, determining that the interaction detection result is no triggered operation on the virtual object.
[0067] Preset posture information can be user-defined posture information for the operating parts, used to achieve specific functions, such as obtaining user requirement information.
[0068] Both the first posture information and the preset posture information can include at least one key point of the operating part. The matching operation between the first posture information and the preset posture information can include: matching at least one key point in the first posture information with at least one key point in the preset posture information.
[0069] Alternatively, at least one key point in the first posture information is connected in a preset order to form a first polygon, and at least one key point in the preset posture information is connected in a preset order to form a second polygon; the shapes of the first polygon and the second polygon are matched.
[0070] In the embodiments of this disclosure, the operation of virtual objects in the virtual world is triggered in a non-contact manner through the first posture information of the operation part, so that users can interact with virtual objects in the virtual world in a simple and convenient way and trigger subsequent acquisition of user demand information.
[0071] According to embodiments of this disclosure, the pose information of the operating part includes first pose information characterizing the pose of the operating part and first position information of the operating part in the real world.
[0072] For operation S220, determining the interaction detection result for the virtual object based on pose information further includes: determining the position detection result based on the first position information and the second position information of the virtual object, wherein the position detection result represents whether the operation part in the real world is in contact with the virtual object in the virtual world, and the second position information represents the position information of the virtual object in the real world; if the position detection result represents that the operation part is in contact with the virtual object and the first pose information matches the preset pose information, the interaction detection result is determined to be a trigger operation on the virtual object.
[0073] Although virtual objects and virtual images exist only in the virtual world, their display position in the virtual world is in the same coordinate system as that of smart devices, such as being in the same world coordinate system.
[0074] For example, if there are no other items on item A in the real world, and the user wears a smart device, they can see item A and a virtual object B above item A. At this time, regardless of whether item A can be viewed in the smart device, with the origin of the fixed world coordinate system, the coordinates of item A in the world coordinate system are (x1, y1, z1). Similarly, even if virtual object B cannot be viewed in the real world, the coordinates of virtual object B in the world coordinate system are (x1, y1, z2).
[0075] In one specific embodiment, the second location information of the virtual object can be the location information of key points in the virtual object. For example, key points can be the model center point, the highest point, etc. of the virtual object; or, key points can also be target parts of the virtual object, such as the top of the virtual object's head, ears, face, etc. Determining the location detection result based on the first location information and the second location information includes: determining the location detection result based on whether the first location information and the second location information are the same.
[0076] In another specific embodiment, the second position information of the virtual object may be the three-dimensional region range information in which the virtual object is located. Determining the position detection result based on the first position information and the second position information includes: determining the position detection result based on whether the first position information falls within the three-dimensional region range information.
[0077] The position detection result indicates that the operating part is in contact with the virtual object and the first posture information matches the preset posture information, indicating that a contact-based interaction has been triggered between the user and the virtual object. For example, a user can pinch the face of a virtual object to trigger an operation on the virtual object in order to subsequently obtain user request information.
[0078] The embodiments of this disclosure can determine whether there is contact between the operating part and the virtual object based on the first position information and the second position information. At the same time, based on the first posture information and the preset posture information of the operating part, it can determine whether to trigger the operation on the operating object through the first posture information, thereby supporting the user to trigger the interaction with the virtual object through contact interaction, so that the user can have a more realistic interaction with the virtual object and improve the user experience.
[0079] According to embodiments of this disclosure, when the position detection result indicates that the operating part is in contact with the virtual object and the first posture information matches the preset posture information, the user's operating part is displayed in the virtual world, so that the user can watch the interaction between the user and the virtual object in the virtual world in real time through a smart device.
[0080] Figure 3 schematically illustrates a scenario of a method for determining interaction detection results for the virtual object according to an embodiment of the present disclosure.
[0081] As shown in Figure 3, the pose information 301 of the operating part includes first pose information 302 and first position information 303, and the position information of the virtual object in the real space is second position information 304. Based on the first position information 303 of the operating part and the second position information 304 of the virtual object, the position detection result 305 can be determined.
[0082] In operation S310, it is determined whether the first attitude information matches the preset attitude information.
[0083] When operating S320, determine whether the operating part is in contact with the virtual object.
[0084] When operating S330, an operation on a virtual object is triggered.
[0085] No operation on the virtual object was triggered when operating S340.
[0086] In one embodiment, for operation S310, if it is determined that the first posture information matches the preset posture information, operation S330 can be directly entered to trigger the operation on the virtual object; if it is determined that the first posture information does not match the preset posture information, operation S340 can be directly entered without triggering the operation on the virtual object.
[0087] In another embodiment, for operation S310, if it is determined that the first posture information matches the preset posture information, and operation S320 determines that the operation part is in contact with the virtual object, operation S330 is entered to trigger the operation on the virtual object; otherwise, operation S340 is entered without triggering the operation on the virtual object.
[0088] Furthermore, in operation S320, if it is determined that the operation part is not in contact with the virtual object, operation S340 is directly entered without triggering the operation on the virtual object.
[0089] According to embodiments of this disclosure, smart devices can collect voice data input by users. However, smart devices are currently unable to recognize the voice content of the collected voice data. Therefore, it is necessary to use the voice input function of virtual objects to recognize the collected voice data.
[0090] In one specific embodiment, in response to detecting a user's operation on a virtual object, the virtual object's voice input function is activated. Then, after activating the virtual object's voice input function, the virtual object's application software can acquire voice data collected through a smart device; and determine user demand information based on the voice data.
[0091] According to embodiments of this disclosure, voice data can be converted into textual user request information using Automatic Speech Recognition (ASR) technology.
[0092] In the embodiments disclosed herein, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of data (e.g., including but not limited to user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard user personal information security and network security.
[0093] For example, in an embodiment of this disclosure, before collecting voice data, the application software of the virtual object and virtual screen will request voice permission from the user, and collect voice data after obtaining the user's authorization or consent.
[0094] According to embodiments of this disclosure, if no sound wave exceeding a certain threshold is detected for a preset duration, the acquisition of voice data is stopped, and the input voice data is automatically converted into user request information in text form. For example, the acquisition of voice data may stop if no sound wave is detected for 0.5 seconds.
[0095] Alternatively, in response to detecting a user's preset action on a virtual object, the system stops collecting voice data and automatically converts the input voice data into text format representing the user's request. For example, the user manually clicks on a virtual object to start collecting voice data, and clicks on the virtual object again to stop collecting voice data.
[0096] According to embodiments of this disclosure, after determining user request information, the user request information, converted from text, can be displayed to the user to avoid affecting the generation of subsequent response information. For example, the user request information can be displayed as text in an area associated with a virtual object.
[0097] In the embodiments of this disclosure, the virtual object belonging to the same application software as the virtual screen also supports voice interaction. The user's needs information is determined through voice interaction between the user and the virtual object, so that the virtual object can obtain the user's needs information in a more intelligent and more sales guide-like manner, providing the user with a more intelligent and user-friendly virtual object.
[0098] In related technologies, due to the limited screen size of mobile or PC devices, virtual objects only appear on specific pages to solve a specific problem, such as digital customer service avatars. In online shopping scenarios, displaying virtual objects on mobile or PC devices essentially means displaying virtual objects on a flat surface, which would obscure a large portion of the shopping page, thus impacting user experience and making it difficult for virtual objects to be integrated throughout the entire user shopping experience.
[0099] Conversely, embodiments of this disclosure display virtual images and virtual objects in a three-dimensional virtual space. In this case, the virtual objects can accompany the user throughout the entire shopping process and provide services to them at any time, just like a "shopping guide" in an offline shopping scenario.
[0100] According to embodiments of this disclosure, the method further includes: determining second position information of a virtual object based on third position information of a virtual image; or, acquiring pose information of a smart device, wherein the pose information of the smart device includes fourth position information and second posture information; determining the second position information based on the fourth position information, the second posture information, and the user's arm span distance; the user's arm span distance characterizes the farthest distance when the user extends their arm.
[0101] According to embodiments of this disclosure, a smart device can also simultaneously display virtual images projected onto multiple application software programs, i.e., a first display state. For example, if a user is running multiple application software programs simultaneously, the smart device can simultaneously display the virtual images of each of the multiple application software programs. If the application software to which the multiple virtual images belong includes virtual objects, the virtual objects can be displayed simultaneously. For example, when the smart device is in a semi-immersive mode, the virtual images can be in the first display state.
[0102] According to embodiments of this disclosure, a smart device may also display only a virtual image projected onto a single application, i.e., a second display state. For virtual objects belonging to the same application as the virtual image of a particular application, the smart device may also display the virtual object simultaneously when the virtual image is in the second display state. Conversely, if no virtual object exists, only the virtual image is displayed. For example, when the smart device is in a fully immersive mode, the virtual image can be in the second display state, and the virtual space at this time includes only a virtual image and a virtual object belonging to the same application.
[0103] After opening the virtual screen application, virtual screens and virtual objects can be displayed simultaneously in the virtual space; virtual objects exist in the virtual space when the application is running or running in the background.
[0104] In one specific embodiment, the positions of the virtual object and the virtual screen can be bound together. Therefore, the second position information of the virtual object is determined based on the third position information of the virtual screen. Specifically, based on a predefined positional relationship, the second position information of the virtual object is determined according to the third position information of the virtual screen in the world coordinate system. It should be noted that the third position information refers to the position information of the virtual screen in the world coordinate system.
[0105] Specifically, the predefined positional relationship can be: the virtual object is within a preset distance range near the virtual screen, so that the virtual object does not obscure the virtual screen; or, the virtual object is located in the lower right corner of the virtual screen. In this embodiment, when the user drags the virtual screen, the virtual object can move along with the virtual screen.
[0106] For example, in the first display state, users wearing smart devices can view virtual images in the virtual world, as well as virtual objects in the lower right corner of the virtual images. That is, from the user's perspective, the virtual objects are in a certain posture in the virtual world.
[0107] Based on the third position information of the virtual image, determine the second position information of the virtual object.
[0108] In another specific embodiment, the position of the virtual object can be related to the pose information of the smart device. The pose information of the smart device includes fourth position information and second pose information, the second pose information being, for example, azimuth information.
[0109] In this embodiment, the second posture information of the smart device can characterize the user's perspective when wearing the smart device. For example, when the smart device is tilted downwards, it indicates that the user may be in a downward-looking state.
[0110] To ensure users can interact with virtual objects at any time, the virtual objects need to be positioned where the user can see them, i.e., within the user's field of vision. Furthermore, if the virtual object is too close, it may startle the user; if it's too far, the user may not be able to "touch" it, thus affecting the interactive experience. Therefore, the position of the virtual object is not only related to the smart device's posture but also to the user's arm span. The user's arm span can be the farthest distance a user can travel with their arm outstretched, calculated based on expert experience.
[0111] Specifically, determining the second position information based on the fourth position information, the second posture information, and the user's arm span distance includes: determining the user's field of vision based on the second posture information; and within the aforementioned field of vision, determining the second position information of the virtual object based on the fourth position information and the user's arm span distance.
[0112] For example, the direction offset by 5° from the center axis of the field of view can be taken as the field of view direction. Taking the fourth position information as the starting point, the position point of the virtual object is taken as the distance from the starting point in the field of view direction that is less than or equal to the user's arm span distance, thereby obtaining the second position information.
[0113] According to embodiments of this disclosure, when a user switches the spatial anchor point of the virtual screen, the second location information of the virtual object can be repositioned, and the display position of the virtual object can be adjusted based on the switched spatial anchor point, so that the user can always find the virtual object in a relatively comfortable position during the shopping experience. The second location information of the virtual object can be repositioned based on the third location information of the virtual screen or the pose information of the smart device. Here, the spatial anchor point can be understood as the real scene or virtual scene that carries the virtual screen. For example, the real scene can be a reproduction of a bedroom or living room in reality, and the virtual scene can be a modeled bedroom or living room.
[0114] The embodiments of this disclosure determine the display state of the virtual screen and adjust the display position of the virtual object based on the display state, so that in either the first or second display state, the user can find the virtual object in a relatively comfortable position throughout the entire shopping experience, thereby improving the user experience.
[0115] According to embodiments of this disclosure, generating response information that matches user demand information based on user demand information includes: determining user intent based on user demand information; and generating response information that matches user demand information based on user intent.
[0116] Determining user intent based on user demand information can include: pre-setting multiple user intents and categorizing them; using a classification algorithm to classify the user demand information; and determining the user intent based on the classification results. For example, for user-inputted demand information (query), the probability of each intent is calculated using a classification model, and finally, the user intent matching the user demand information is determined.
[0117] Alternatively, various intent recognition algorithms can be used to determine user intent. Intent recognition algorithms include, but are not limited to, at least one of the following: machine learning algorithms, such as Support Vector Machines (SVM), and deep learning algorithms, such as Long Short-Term Memory (LSTM), Bi-RNN, and Bi-LSTM-CRF.
[0118] For example, in online shopping scenarios, user intent may include: recommending products, providing shopping suggestions, supplementing product category background information, and voice communication.
[0119] Generating response information that matches the user's needs based on the user's intent includes: determining a question-and-answer database that matches the user's intent; searching the question-and-answer database for the question-and-answer pair most similar to the user's needs, and using the answer of that question-and-answer pair as the response information. In some embodiments, the response information can also be determined by combining the user's needs information and historical needs information.
[0120] The embodiments of this disclosure determine user intent based on user demand information and generate response information that matches the user demand information according to the user intent. This can lead to a strategy-based distribution of response information generation based on user intent, which not only makes the generated response information more in line with the user's true intent and needs, but also enables the rapid generation of response information based on user intent.
[0121] Because application software involves numerous functions, virtual objects need to interact with various virtual screens and provide diverse response information to users based on these screens. Therefore, Large Language Models (LLMs) can be integrated with virtual objects, enabling them to provide comprehensive response information to users in real time.
[0122] For example, a large language model can be used to determine user intent based on user demand information. Alternatively, historical user demand information can be used as contextual information, and together with the current user demand information, it can be used as input to the large language model to determine user intent. Historical demand information can refer to the demand information input by the user through interaction with the virtual object after its appearance.
[0123] According to embodiments of this disclosure, determining user intent based on user demand information includes: obtaining historical demand information and the current screen content of a virtual screen, wherein the current screen content includes at least one screen content corresponding to each of multiple shopping functions in the shopping platform; and using a large language model to determine user intent based on user demand information, historical demand information, and the current screen content.
[0124] In the embodiments of this disclosure, application software typically includes various functional pages, each function usually includes multiple display pages, and may even include a specific page tailored to the user's individual needs. Therefore, when these pages are projected as virtual images via a smart device, the virtual objects are accompanied by various virtual images.
[0125] Based on the description of the above embodiments, users can see virtual objects within a preset distance range in the lower right corner of the virtual screen, and the presence of these virtual objects will not obstruct the virtual screen; alternatively, users can also see virtual objects at a distance of arm's length from the center of their field of vision. This can be understood as the virtual object accompanying the user in a comfortable position throughout the process of using the application software.
[0126] User demand information includes demand information for virtual screens, and virtual screens can be projections of pages of various formats. Therefore, embodiments of this disclosure take into account the current screen content of the virtual screen, that is, to locate the user's position in the entire process, in order to provide more accurate response information.
[0127] Specifically, user demand information, historical demand information, and the current screen content can be used as inputs to the large language model, such as a prompt, and the large language model can output the user's intent. When constructing the input to the large language model, prompts such as "output user intent" can also be added.
[0128] According to embodiments of this disclosure, historical demand information can be obtained through the historical interaction logs between the user and the virtual object. Historical demand information refers to the demand information input by the user through interaction with the virtual object after its appearance, and this information can be displayed to the user in the form of a dialogue.
[0129] According to embodiments of this disclosure, since the virtual image is projected through a smart device and the display page before projection is generated by application software, the application software page before projection can be used as the current image content of the virtual image.
[0130] For example, taking an application software as a shopping platform, the virtual object can be seen as a "smart shopping guide" that accompanies the user throughout the entire shopping process. Users can view the shopping platform's homepage (virtual screen 1) through their smart devices. When a user opens the shopping platform, the virtual object can appear before them, such as welcoming them with an opening animation. After clicking on a function within the shopping page, the user can view the corresponding functional area through their smart device, such as the discount section for certain products (virtual screen 2). Virtual screens 1 and 2, along with the virtual object, all belong to the shopping platform.
[0131] For virtual screen 1, the user's request information could be "Where is function XX? I can't find it," while historical request information could be "I've never used this shopping platform before, how do I use it?" and the current screen content, such as the shopping homepage. Therefore, the large language model can determine the user's intent as a function index based on the user's request information, historical request information, and the current screen content. Alternatively, the user's request information could be "What's fun to do?", combined with the current screen content "shopping homepage," which can determine the user's intent as a function recommendation or a product recommendation.
[0132] For virtual screen 2, the user's request information could be "What are some recommendations?", and combined with the current screen content "discount area product page", it can be determined that the user's intention is product recommendation.
[0133] The embodiments of this disclosure utilize a large language model, combining user demand information, historical demand information, and current screen content, to generate more accurate response information from three perspectives: real-time demand, historical demand, and current visual perception, thereby improving user experience.
[0134] According to embodiments of this disclosure, generating response information that matches user needs based on user intent includes: determining a response strategy that matches user intent; and generating response information based on user needs and the response strategy using a large language model.
[0135] Multiple user intents can be matched with multiple response strategies, and these response strategies can include different question-and-answer databases. The matching relationship between user intents and response strategies can be pre-defined. Therefore, after determining the user intent based on user needs information, a response strategy matching that user intent can be determined based on the pre-defined matching relationship.
[0136] Alternatively, the response strategy can be a pre-defined expert rule or response template. Therefore, after determining the user's intent based on user needs information, an expert rule or response template matching that intent can be determined based on pre-defined matching relationships.
[0137] In embodiments of this disclosure, the input to the large language model can be generated based on user needs information and response strategies, such as "generating an answer that meets the user's needs information based on the response strategy," so that the large language model can output response information.
[0138] In embodiments of this disclosure, the large language model can invoke a toolbox corresponding to the response strategy to generate response information that matches the user's needs.
[0139] The embodiments of this disclosure utilize user intent to determine the response strategy, which can achieve strategy diversion for generating response information, while improving the speed and accuracy of response information generation.
[0140] According to embodiments of this disclosure, the method further includes: using a large language model to determine response information based on user demand information, historical demand information, and response strategies.
[0141] According to embodiments of this disclosure, a large language model is used to determine response information based on user demand information, historical demand information, the current content of the virtual screen, and a response strategy. Similar to determining user intent, since the virtual screen can be a projection of pages with various formats, taking into account the current content of the virtual screen can locate the user's current position throughout the entire process, thereby providing more accurate response information.
[0142] Figure 4 schematically illustrates a flowchart of a method for determining response information according to an embodiment of the present disclosure.
[0143] As shown in Figure 4, in embodiment 400, the large language model M10 can be used to determine the user's intent and generate response information. In other embodiments, only the large language model can be used to determine the user's intent or generate response information, and there is no limitation thereto.
[0144] The large language model M10 can integrate multiple model components to perform various tasks. In the embodiments of this disclosure, the large language model M10 includes at least a first model component M11 and a second model component M12, which are used to determine user intent and generate response information, respectively.
[0145] User request information 401, historical request information 402, and the current screen content 403 of the virtual screen are input into the first model component M11 of the large language model M10, and the user intent is output. After determining the response strategy 404 based on the user intent, the response strategy 404 can be input into the second model component M12 of the large language model M10. The second model component M12 can generate response information 405 according to the response strategy 404, user request information 401, historical request information 402, and the current screen content 403 of the virtual screen.
[0146] According to embodiments of this disclosure, displaying response information includes: displaying the response information in text form to a user through a display window in a virtual world, wherein the position information of the display window is determined based on a second position information of the virtual object; or, converting the response information into speech data and generating a lip-shape change sequence matching the speech data; generating a virtual object speech response animation based on the lip-shape change sequence and model parameters of the virtual object; and outputting the speech data through a smart device and synchronously displaying the virtual object speech response animation in the virtual world.
[0147] According to embodiments of this disclosure, the display window in the virtual world can be a display area associated with a virtual object. This display window can be in the style of a chat dialog box, where the virtual object can display user request information and matching response information to the user in the form of chat bubbles. Furthermore, this display area also supports further user interaction through clicking or other means.
[0148] There is a positional relationship between the display window and the virtual object; for example, the display window is usually positioned above and behind the virtual object.
[0149] In another specific embodiment, the virtual object can also simulate a real dialogue scenario by outputting voice messages to the user.
[0150] Specifically, after generating the reply information, text-to-speech technology can be used to convert the reply information into speech data, such as deep learning-based text-to-speech (TTS).
[0151] Simultaneously, a lip-shape sequence can be generated based on the response information. The speed of two lip-shape changes in the sequence matches the speed of two syllable changes in the speech data, avoiding audio-visual asynchrony. To create a realistic interaction scenario with the virtual object, the virtual object can also "speak" the response information. Specifically, the converted speech data can be output through the voice playback unit of the smart device, while the smart device displays the virtual object's voice response animation, thus showing the user a virtual object that is "speaking" synchronously.
[0152] In embodiments of this disclosure, lip response animation can be generated based on a lip shape change sequence and model lip shape parameters; subsequently, the lip response animation can be fused with other model parameters of the virtual object to generate a virtual object speech response animation.
[0153] In the embodiments of this disclosure, since the response information is generated based on the user's operation on the virtual object, the response information can be fed back to the user through the display method associated with the virtual object, so that the interaction process between the user and the virtual object is similar to the dialogue method of a real shopping guide, thereby improving the user experience.
[0154] Figure 5 schematically illustrates a flowchart of displaying virtual object animations and reply information according to another embodiment of the present disclosure.
[0155] As shown in Figure 5, the method 500 for displaying virtual object animations and replying to information includes operations S510 to S530.
[0156] When operating S510, the user's emotional information is determined based on user demand information and / or historical demand information.
[0157] In operation S520, a virtual object display animation that matches the emotional information is determined, wherein the virtual object display animation indicates changes in the virtual object's display actions and / or changes in its display expressions.
[0158] While operating the S530, both reply messages and virtual object display animations are shown.
[0159] According to embodiments of this disclosure, in order to enable virtual objects to provide more human-like feedback to users, virtual objects can also have corresponding action changes and facial expression changes while displaying feedback information.
[0160] In one specific embodiment, a virtual object display animation, such as a smiling or waving animation, can be pre-set, and a reply message and the aforementioned virtual object display animation can be displayed to the user at the same time.
[0161] In another specific embodiment, the user's emotional information can be determined based on user demand information and / or historical demand information, and a virtual object display animation that matches the emotional information can be determined, so that the generated virtual object display animation is more in line with the user's current emotional needs.
[0162] Specifically, emotion recognition technology can be used to determine a user's emotional information from their needs information, or it can be combined with historical needs information to determine the user's emotional information. There is a pre-defined matching relationship between emotional information and virtual object display animations. For example, when the emotional information is "happy," the virtual object display animation could be jumping or laughing. When the emotional information is "unhappy," the virtual object display animation could be smiling.
[0163] The embodiments of this disclosure determine the user's current emotional information by using user demand information and / or historical demand information, and match appropriate virtual objects to display animations and response information together to present to the user, which can make the virtual objects more human-like and meet the user's emotional needs.
[0164] When the embodiments of this disclosure are applied to an online shopping platform, the virtual object can act as an intelligent shopping guide to accompany the user throughout the entire shopping process. Through the interaction between the user and the intelligent shopping guide, coupled with the underlying large model and toolbox for problem analysis and response, the user's online shopping mode is upgraded from only being able to purchase low-threshold, highly standardized products on the platform to having the platform's experienced intelligent shopping guide analyze and break down their needs, thereby supporting the user's purchase of non-standardized, high-threshold products, and providing assistance to the user in making shopping decisions at any time.
[0165] Furthermore, virtual objects can interact with users as intelligent shopping guides based on spatial computing logic. Unlike ordinary 2D rendering, virtual objects based on spatial computing carry depth information, creating the feeling that the virtual objects are truly present around the user.
[0166] Figure 6 schematically illustrates a block diagram of a virtual object-based interactive device according to an embodiment of the present disclosure.
[0167] As shown in Figure 6, the virtual object-based interactive device 600 includes a first acquisition module 610, an interaction detection module 620, a second acquisition module 630, a generation module 640, and a display module 650.
[0168] The first acquisition module 610 is used to acquire the position and pose information of the user's operating part in the real world. The position and pose information is determined based on the image captured by the smart device. The smart device is also used to display a virtual world to the user. The virtual world includes virtual images and virtual objects belonging to the same application software as the virtual images.
[0169] The interaction detection module 620 is used to determine the interaction detection results for virtual objects based on pose information.
[0170] The second acquisition module 630 is used to acquire user demand information in response to the determination of the interaction detection result characterization triggering the operation on the virtual object, wherein the user demand information includes demand information for the virtual screen.
[0171] The generation module 640 is used to generate response information that matches the user's needs information.
[0172] Display module 650 is used to display reply information.
[0173] According to embodiments of this disclosure, the interaction detection module 620 includes a first pose detection unit and a second pose detection unit. The pose information includes the first pose information.
[0174] The first posture detection unit is used to determine that the interaction detection result is to trigger an operation on the virtual object when the first posture information matches the preset posture information.
[0175] The second posture detection unit determines that the interaction detection result is no operation on the virtual object if it finds that the first posture information does not match the preset posture information.
[0176] According to embodiments of this disclosure, the interaction detection module 620 further includes a position detection unit and an interaction detection unit. The pose information also includes first position information.
[0177] The position detection unit is used to determine the position detection result based on the first position information and the second position information of the virtual object. The position detection result indicates whether the operating part in the real world is in contact with the virtual object in the virtual world, and the second position information indicates the position information of the virtual object in the real world.
[0178] The interaction detection unit is used to determine that the interaction detection result is a trigger for an operation on the virtual object when the position detection result indicates that the operation part is in contact with the virtual object and the first posture information matches the preset posture information.
[0179] According to embodiments of this disclosure, the virtual object-based interactive device 600 further includes: a first position determination module and / or a second position determination module.
[0180] The first position determination module is used to determine the second position information of the virtual object based on the third position information of the virtual image.
[0181] The second position determination module is used to acquire the pose information of the smart device, wherein the pose information of the smart device includes fourth position information and second posture information; the second position information is determined based on the fourth position information, the second posture information and the user's arm span distance; the user's arm span distance represents the farthest distance when the user extends their arm.
[0182] According to embodiments of this disclosure, the generation module 620 includes an intent determination unit and a generation unit. The intent determination unit is used to determine a user intent based on user demand information. The generation unit is used to generate response information that matches the user demand information based on the user intent.
[0183] According to embodiments of this disclosure, the intent determination unit includes an acquisition subunit and a first determination subunit.
[0184] The acquisition sub-unit is used to acquire historical demand information and the current screen content of the virtual screen. The current screen content includes at least one screen content corresponding to each of the multiple shopping functions in the shopping platform.
[0185] The first determining subunit is used to determine the user's intent by utilizing a large language model based on user demand information, historical demand information, and the current screen content.
[0186] According to embodiments of this disclosure, the generation unit includes a second determining subunit and a generation subunit. The second determining subunit is used to determine a response strategy that matches the user's intent. The generation subunit is used to generate response information based on user needs information and the response strategy, utilizing a large language model.
[0187] According to embodiments of this disclosure, the display module 630 includes a first display unit or a second display unit. The first display unit is used to display text-based response information to the user through a display window in a virtual world, wherein the position information of the display window is determined based on second position information of the virtual object. The second display unit is used to convert the response information into speech data and generate a lip-sync sequence matching the speech data; generate a virtual object speech response animation based on the lip-sync sequence and model parameters of the virtual object; output the speech data through a smart device and synchronously display the virtual object speech response animation in the virtual world.
[0188] According to embodiments of this disclosure, the virtual object-based interactive device 600 includes an emotion determination module, an animation determination module, and an animation display module.
[0189] The emotion determination module is used to determine the user's emotional information based on user needs and / or historical needs information. The animation determination module is used to determine virtual object display animations that match the emotional information, where the virtual object display animations indicate changes in the virtual object's display actions and / or facial expressions. The animation display module is used to simultaneously display the response information and the virtual object display animations.
[0190] Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure, or at least part of the functions of any one or more of them, can be implemented in one module. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be implemented by dividing them into multiple modules. Any one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as hardware circuitry, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Array (PLA), a System-on-Chip, a System-on-a-Substrate, a System-on-Package, an Application-Specific Integrated Circuit (ASIC), or implemented in hardware or firmware by any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, one or more of the modules, submodules, units, and subunits according to embodiments of the present disclosure can be at least partially implemented as computer program modules, which, when run, can perform corresponding functions.
[0191] For example, any plurality of the acquisition module 610, generation module 620, and display module 630 may be combined into one module / unit / subunit, or any one of these modules / units / subunits may be split into multiple modules / units / subunits. Alternatively, at least part of the functionality of one or more of these modules / units / subunits may be combined with at least part of the functionality of other modules / units / subunits and implemented in one module / unit / subunit. According to embodiments of the present disclosure, at least one of the acquisition module 610, generation module 620, and display module 630 may be at least partially implemented as hardware circuitry, such as a field-programmable gate array (FPGA), a programmable logic array (PLA), a system-on-a-chip, a system-on-a-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or any other reasonable means of integrating or packaging circuitry, or implemented in software, hardware, or firmware, or in any suitable combination of any of these three implementation methods. Alternatively, at least one of the acquisition module 610, the generation module 620, and the display module 630 may be implemented at least partially as a computer program module that can perform corresponding functions when the computer program module is run.
[0192] It should be noted that the description of the apparatus portion in the embodiments of this disclosure is specifically referred to in the method portion, and will not be repeated here.
[0193] Figure 7 schematically illustrates a block diagram of an electronic device suitable for implementing a virtual object-based interaction method and apparatus according to embodiments of the present disclosure. The electronic device shown in Figure 7 is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present disclosure.
[0194] As shown in FIG. 7, an electronic device 700 according to an embodiment of the present disclosure includes a processor 701, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 702 or a program loaded from a storage portion 708 into a random access memory (RAM) 703. The processor 701 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or an associated chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 701 may also include onboard memory for caching purposes. The processor 701 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0195] RAM 703 stores various programs and data required for the operation of electronic device 700. Processor 701, ROM 702, and RAM 703 are interconnected via bus 704. Processor 701 performs various operations of the method flow according to embodiments of the present disclosure by executing programs in ROM 702 and / or RAM 703. It should be noted that the programs may also be stored in one or more memories other than ROM 702 and RAM 703. Processor 701 may also perform various operations of the method flow according to embodiments of the present disclosure by executing programs stored in said one or more memories.
[0196] According to embodiments of this disclosure, the electronic device 700 may further include an input / output (I / O) interface 705, which is also connected to a bus 704. The electronic device 700 may also include one or more of the following components connected to the input / output (I / O) interface 705: an input section 706 including a keyboard, mouse, etc.; an output section 707 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN card, modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output (I / O) interface 705 as needed. A removable medium 711, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive 710 as needed so that computer programs read from it can be installed into the storage section 708 as needed.
[0197] According to embodiments of this disclosure, the method flow according to embodiments of this disclosure can be implemented as a computer software program. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing program code for performing the methods shown in the flowchart. In such embodiments, the computer program can be downloaded and installed from a network via communication section 709, and / or installed from removable medium 711. When the computer program is executed by processor 701, it performs the functions defined in the system of embodiments of this disclosure. According to embodiments of this disclosure, the systems, devices, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0198] This disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or it may exist independently and not assembled into the device / apparatus / system. The computer-readable storage medium carries one or more programs that, when executed, implement the method according to the embodiments of this disclosure.
[0199] According to embodiments of this disclosure, the computer-readable storage medium can be a non-volatile computer-readable storage medium. Examples include, but are not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0200] For example, according to embodiments of this disclosure, a computer-readable storage medium may include the ROM 702 and / or RAM 703 described above and / or one or more memories other than ROM 702 and RAM 703.
[0201] Embodiments of this disclosure also include a computer program product comprising a computer program containing program code for performing the methods provided in the embodiments of this disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the virtual object-based interactive method provided in the embodiments of this disclosure.
[0202] When the computer program is executed by the processor 701, it performs the functions defined in the system / apparatus of this disclosure embodiments. According to embodiments of this disclosure, the systems, apparatuses, modules, units, etc., described above can be implemented by computer program modules.
[0203] In one embodiment, the computer program may rely on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may also be transmitted and distributed in the form of signals over a network medium, and may be downloaded and installed via the communication section 709, and / or installed from a removable medium 711. The program code contained in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination thereof.
[0204] According to embodiments of this disclosure, program code for executing the computer programs provided in embodiments of this disclosure can be written in any combination of one or more programming languages. Specifically, these computational programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include, but are not limited to, languages such as Java, C++, Python, "C", or similar programming languages. The program code can execute entirely on a user's computing device, partially on a user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0205] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions. Those skilled in the art will understand that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations are not explicitly described in the present disclosure. In particular, the features described in the various embodiments of this disclosure may be combined and / or combined in various ways without departing from the spirit and teachings of this disclosure. All such combinations and / or combinations fall within the scope of this disclosure.
[0206] The embodiments of this disclosure have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of this disclosure. Although various embodiments have been described above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of this disclosure, and all such substitutions and modifications should fall within the scope of this disclosure.
Claims
1. A virtual object based interaction method, comprising: obtaining pose information of an operating part of a user in a real world, wherein the pose information is determined according to an image captured by a smart device, the smart device is further configured to show a virtual world to the user, the virtual world comprises a virtual picture and a virtual object belonging to a same application software as the virtual picture; determining an interaction detection result for the virtual object based on the pose information; in response to determining that the interaction detection result represents triggering an operation on the virtual object, obtaining user demand information, the user demand information comprising demand information for the virtual picture; generating reply information matching the user demand information based on the user demand information; and showing the reply information.
2. The method of claim 1, wherein, The pose information comprises first attitude information; and the determining the interaction detection result for the virtual object based on the pose information comprises: in a case where the first attitude information matches preset attitude information, determining that the interaction detection result is triggering the operation on the virtual object; and in a case where the first attitude information does not match the preset attitude information, determining that the interaction detection result is not triggering the operation on the virtual object.
3. The method of claim 2, wherein, The pose information further comprises first position information, and the determining the interaction detection result for the virtual object based on the pose information further comprises: determining a position detection result according to the first position information and second position information of the virtual object, wherein the position detection result represents whether the operating part in the real world contacts the virtual object in the virtual world, and the second position information represents position information of the virtual object in the real world; and in a case where the position detection result represents that the operating part contacts the virtual object and the first attitude information matches the preset attitude information, determining that the interaction detection result is triggering the operation on the virtual object. 4.The method of claim 3, further comprising: determining the second position information of the virtual object according to third position information of the virtual picture; or obtaining pose information of the smart device, wherein the pose information of the smart device comprises fourth position information and second attitude information; and determining the second position information according to the fourth position information, the second attitude information and a user arm span distance, wherein the user arm span distance represents a farthest distance when the user spreads arms. The generating the reply information matching the user demand information based on the user demand information comprises:
5. The method of claim 1, wherein, determining a user intention based on the user demand information; and generating the reply information matching the user demand information according to the user intention. The determining the user intention based on the user demand information comprises:
6. The method of claim 5, wherein, obtaining historical demand information and current picture content of the virtual picture, wherein the current picture content comprises at least one picture content corresponding to each of a plurality of shopping functions in a shopping platform. The user intention is determined by using a large language model according to the user demand information, historical demand information, and the current picture content.
7. The method of claim 5, wherein, The reply information matched with the user demand information is generated according to the user intention, including: A reply strategy matched with the user intention is determined; The reply information is generated based on the user demand information and the reply strategy by using the large language model.
8. The method of claim 1, wherein, The reply information is displayed, including: The reply information in the form of text is displayed to the user through a display window in the virtual world, wherein the position information of the display window is determined according to the second position information of the virtual object; or The reply information is converted into voice data, and a lip change sequence matched with the voice data is generated; a virtual object voice reply animation is generated according to the lip change sequence and model parameters of the virtual object; the voice data is output through the intelligent device, and the virtual object voice reply animation is displayed in the virtual world at the same time.
9. The method of claim 1 or 8, further comprising: determining emotional information of the user according to the user demand information and / or historical demand information; determining a virtual object display animation matched with the emotional information, wherein the virtual object display animation indicates display action changes and / or display expression changes of the virtual object; displaying the reply information and the virtual object display animation at the same time.
10. An interactive device based on a virtual object, comprising: a first acquisition module configured to acquire pose information of an operation part of a user in a real world, wherein the pose information is determined according to an image captured by an intelligent device, the intelligent device is further configured to display a virtual world to the user, the virtual world includes a virtual picture and a virtual object belonging to the same application software as the virtual picture; an interaction detection module configured to determine an interaction detection result for the virtual object based on the pose information; a second acquisition module configured to acquire user demand information in response to determining that the interaction detection result represents triggering an operation on the virtual object, wherein the user demand information includes demand information for the virtual picture; a generation module configured to generate reply information matched with the user demand information based on the user demand information; and a display module configured to display the reply information.
11. An electronic device, comprising: one or more processors; a memory configured to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 9.
12. A computer-readable storage medium having stored thereon executable instructions that, when executed by a processor, cause the processor to implement the method of any one of claims 1 to 9.
13. A computer program product comprising a computer program that, when executed by a processor, implements the method of any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-modal interactive processing method and system based on virtual person
CN107894833A
Virtual character control method and device, apparatus and readable storage medium
CN110308792A
AR scene interaction method and device, electronic equipment and storage medium
CN112148189A
Virtual character interaction strategy determination method and device
CN116009692A
Interaction method and cloud server
CN116225234A