Virtual object interaction method, electronic device, and computer-readable storage medium

By using waiting pages, connection progress bars, carousel areas, and guide buttons during virtual human interaction, the anxiety of users during the connection process was resolved, the interactive experience and brand image were improved, and user engagement and commercial value were enhanced.

CN117076798BActive Publication Date: 2026-02-06MOFA (SHANGHAI) INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310572265.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-19
Publication Date
2026-02-06
Estimated Expiration
2043-05-19

AI Technical Summary

Technical Problem

In existing service-oriented virtual human interaction processes, users are prone to anxiety and impatience during the connection establishment process, resulting in a poor interaction experience.

Method used

During the connection process, a waiting page is displayed, a connection progress bar and a carousel area are set, and a guide page and guide button are displayed after the connection is established. The corresponding playback page is displayed according to the user's identifier, and function buttons are provided in the side and bottom navigation areas.

Benefits of technology

It alleviated users' waiting anxiety, improved the interactive experience and smoothness, enhanced user engagement and stickiness, improved brand image and commercial value, and reduced maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117076798B_ABST
    Figure CN117076798B_ABST
Patent Text Reader

Abstract

The application provides a virtual object interaction method, an electronic device and a computer readable storage medium, which are used for realizing a virtual object interaction function, and the method comprises the following steps: in response to an access request of a terminal device, a waiting page is displayed by using the terminal device in the process of establishing a connection between the terminal device and a target server, and the access request is used to indicate a user identifier; after the connection is established, a guide page is displayed by using the terminal device, and the guide page is provided with a guide button; and in response to a click operation on the guide button, a playing page corresponding to the user identifier is displayed by using the terminal device. In the process of establishing the connection, the waiting page is used to alleviate the anxious mood of the user in waiting, and the playing page of the explanation video group is entered only under the active click operation of the user, so that the ritual sense of the virtual object interaction process is improved, and the commercial value of the virtual object is maintained.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of virtual humans, interaction design, artificial intelligence, and in particular to a virtual object interaction method, an electronic device and a computer readable storage medium. BACKGROUND

[0002] Virtual human (virtual human or computer synthesized character) refers to the representation of geometric characteristics and behavior characteristics of a person in a computer-generated space (virtual environment). The main function of a service-type virtual human is to replace a real person to provide services and provide daily companionship. It is a virtualization of a service-type role in reality, including virtual anchors, virtual teachers, virtual customer service representatives, virtual assistants, etc. Its industrial value is mainly to reduce the cost of existing service-type industries and to reduce costs and increase efficiency for the stock market. The fields in which service-type virtual humans (also known as virtual employees) are currently implemented include virtual news broadcasters, virtual anchors, digital tour guides, virtual financial consultants, virtual lawyers, etc.

[0003] The AI video interaction management background can configure the playback logic, content, application style, etc. of the virtual human according to needs, thereby generating an H5 (HTML5) carrier of a virtual human interaction video that can be seen by a user. Since the existing service-type virtual human interaction process requires a series of video clips to be played, a connection needs to be established with a video push service, and in the process of establishing the connection, the user often feels anxious and impatient, and the interaction experience is not good.

[0004] Based on this, the present application provides a virtual object interaction method, an electronic device and a computer readable storage medium to improve the prior art. SUMMARY

[0005] The purpose of the present application is to provide a virtual object interaction method, an electronic device and a computer readable storage medium, which alleviates the anxious mood of the user waiting in the process of establishing a connection, and improves the interaction experience.

[0006] The purpose of the present application is achieved by adopting the following technical solutions:

[0007] In a first aspect, the present application provides a virtual object interaction method for implementing a virtual object interaction function, the method comprising:

[0008] In response to an access request of a terminal device, a waiting page is displayed by the terminal device in the process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0009] After the connection is established, a guide page is displayed by the terminal device, the guide page being provided with a guide button;

[0010] In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by using the terminal device.

[0011] The technical scheme has the beneficial effects that: in the process of establishing a connection, the waiting page is used to alleviate the anxiety of the user in waiting, and the interactive experience is improved. After the connection is established, the guide page and the guide button are used to clearly and intuitively guide the user to enter the play page, and the fluency and comfort of user access are improved. On the other hand, the connection between the terminal device and the target server is established, so that the user can more smoothly watch the explanation video group and more conveniently select the explanation video group, and the user experience is improved. On the other hand, the corresponding play page is displayed according to the user identifier, so that the play page can be personalized, the participation and stickiness of the user are better improved, the interest and attention of the user to the explanation video group are enhanced, and the stickiness of the user to the virtual object interactive mode is improved. On the other hand, the separate guide page and the guide button are set, and the user enters the play page of the explanation video group only under the active click operation of the user. When the user triggers an access request by mistake, the explanation video group will not be automatically played, unnecessary repulsion caused thereby is avoided, that is, a threshold is set for entering the play page, the ritual sense of the virtual object interactive process is improved, and the commercial value of the virtual object is maintained.

[0012] In some possible implementation manners, the waiting page is provided with a connection progress bar and a carousel area.

[0013] The connection progress bar is used to display the connection progress in real time.

[0014] The carousel area is used to carousel a plurality of carousel images, or the carousel area is used to carousel a plurality of carousel videos.

[0015] The beneficial effects of this technical solution are as follows: By incorporating a connection progress bar and a carousel area into the waiting page, users can track the connection progress in real time while waiting for it to be established. The carousel of multiple images or videos further enhances user attention and improves the user experience. Firstly, the progress bar reduces user anxiety by providing real-time updates on the connection's progress. The carousel displays multiple images or videos, offering more information and entertainment, thus increasing user satisfaction and overall experience. Secondly, a well-designed waiting page showcases the company's or product's professionalism and innovation, strengthening brand image and awareness, and earning positive reviews and a better reputation. Thirdly, the carousel's attractive images or videos draw user attention, increasing engagement and interaction, thereby improving marketing effectiveness and conversion rates. Finally, the modular design of the waiting page facilitates maintenance and updates to different sections. Its modular design, high configurability, unified management, and ease of expansion and updates improve the maintainability and scalability of the waiting page, reducing maintenance costs and increasing development efficiency. On the other hand, depending on different application scenarios, the style and content of the connection progress bar and carousel area can be customized to adapt to different user needs and application scenarios, thereby improving the flexibility and adaptability of the application.

[0016] In some possible implementations, the process of obtaining the playback page corresponding to the user identifier includes:

[0017] Obtain the playback type corresponding to the user identifier;

[0018] When the playback type is not played, the initial playback page is used as the playback page corresponding to the user identifier. The initial playback page is used to play multiple explanatory video groups in a preset order.

[0019] When the playback type is playback interruption, the most recently interrupted playback page is used as the playback page corresponding to the user identifier. The interrupted playback page is used to play the explanatory video group corresponding to the interrupted playback page and the explanatory video group in the order following it in the preset order.

[0020] When the playback type is "Playback Ended", the end playback page will be used as the playback page corresponding to the user identifier.

[0021] The beneficial effects of the technical solution are that different playing pages can be provided for the user according to the playing type (i.e., the playing progress type) of the user. When the playing type of the user is not played, the terminal device displays an initial playing page, and the user plays the multiple explanation video groups in a preset order. When the playing type of the user is interrupted playing, the terminal device displays an interrupted playing page corresponding to the last time of interrupted playing, and the user continues to play the explanation video group corresponding to the interrupted playing page and the explanation video groups in an order after the explanation video group in the preset order. When the playing type of the user is end playing, the terminal device displays an end playing page. In this way, the appropriate playing page can be directly reached according to the playing progress of the user, and a better viewing experience can be obtained. By recording the playing state and position of the user, playing the multiple explanation video groups in a preset order, and supporting subsequent viewing after interrupted playing, the user viewing experience, the video viewing time length, and the video playing efficiency can be improved, the use value and competitiveness of the platform can be improved, and the use value and competitiveness of the platform can be improved.

[0022] In some possible implementation manners, the process of obtaining the multiple explanation video groups includes:

[0023] Obtaining device configuration information and network configuration information corresponding to the terminal device;

[0024] Obtaining a definition type corresponding to the device configuration information and the network configuration information;

[0025] Obtaining the multiple explanation video groups corresponding to the definition type.

[0026] The beneficial effects of the technical solution are that the appropriate definition type can be automatically selected according to the configuration information and the network configuration information of the terminal device, and the multiple explanation video groups corresponding to the definition type are obtained. In this way, the user experience can be improved, and the problem of unsmooth video playing caused by poor device configuration or network condition can be avoided. On the one hand, the definition type suitable for the device and the network is automatically selected according to the device configuration information and the network configuration information of the user, so that a better viewing experience can be provided. On the other hand, by selecting the appropriate definition type, the problems such as video lag and slow loading can be avoided, and the viewing quality can be improved. On the other hand, by selecting the appropriate definition type, the data transmission amount can be reduced, and the traffic consumption of the user can be reduced. On the other hand, by providing personalized video viewing experience, the satisfaction and reliability of the user to the product or service can be improved, and the user stickiness can be enhanced. On the other hand, by automatically selecting the definition type, the manual intervention and adjustment can be reduced, and the maintenance cost of the product or service can be reduced.

[0027] In some possible implementation manners, each of the playing pages is provided with a side navigation area and / or a bottom navigation area.

[0028] The side navigation area is provided with one or more of the following function buttons: catalog, subscription, sharing, form, trial, cooperation and external link jump.

[0029] The bottom navigation area is provided with one or more of the following function buttons: pause / resume, catalog, mode switching, hold speaking and asking questions.

[0030] The beneficial effects of the technical scheme are that the side navigation area and / or the bottom navigation area are set for each playing page, and rich function buttons are provided to enable users to have a better viewing experience. The function buttons of the side navigation area include catalog, subscription, sharing, form, trial, cooperation and external link jump, etc., which can meet various needs of users when watching the explanation video group. For example, the user can quickly jump to the video node that the user wants to watch by clicking the catalog button, subscribe to the channel of interest by clicking the subscription button, guide the user to subscribe to the public number, the host, the blogger, etc., share the video with friends by clicking the sharing button, fill in the relevant form by clicking the form button, try the relevant product by clicking the trial button, cooperate with the video provider by clicking the cooperation button, and jump to the relevant webpage by clicking the external link jump button. In addition, the bottom navigation area provides function buttons such as pause / resume, catalog, mode switching, hold speaking and asking questions, which provides more control and interaction for users, and facilitates the user to operate when watching the video. The user can more flexibly control the video playing process, and can ask questions at any time.

[0031] In some possible implementation manners, the bottom navigation area is provided with a question asking button;

[0032] The method further includes:

[0033] In response to a click operation on the question asking button, an explanation video group being played by the terminal device is acquired;

[0034] Based on the explanation video group being played by the terminal device, N preset questions corresponding to the explanation video group are acquired, N being a positive integer;

[0035] The terminal device is used to display the N preset questions in the form of a floating layer;

[0036] In response to a click operation on one of the preset questions, an explanation video group corresponding to the clicked preset question is acquired;

[0037] The terminal device is used to play the explanation video group corresponding to the clicked preset question.

[0038] The technical scheme has the beneficial effects that: by clicking the question button in the bottom navigation area, the user can obtain N preset questions corresponding to the explanation video group being played by the terminal device, and the preset questions are displayed in the form of a floating layer, and the user can click one of the questions to obtain the corresponding explanation video group. In this way, the user can quickly find the information they need, so as to better understand and master the relevant knowledge. On the one hand, the user can quickly obtain the required information and knowledge by clicking the question button, and can quickly understand the key points of the video content through the preset questions, thereby improving the learning efficiency. On the other hand, by providing preset questions and corresponding explanation video groups, the user can more vividly and intuitively understand the video content, thereby improving the interest of learning and improving the user's interest in learning. On the other hand, the user is allowed to click the question button and obtain the corresponding preset questions according to the played explanation video group, thereby improving the user's participation, and the user can more actively participate in learning and exploration. On the other hand, by displaying the preset questions in the form of a floating layer in the terminal device, the user can more conveniently obtain the required information without leaving the current playing page, thereby improving the user experience.

[0039] In some possible implementation manners, each explanation video group is provided with one or more video nodes;

[0040] The N preset questions corresponding to the explanation video group being played by the terminal device are obtained based on the playing progress of the explanation video group being played by the terminal device, and the N preset questions corresponding to the most recent video node are obtained.

[0041] The N preset questions corresponding to the explanation video group being played by the terminal device are obtained based on the playing progress of the explanation video group being played by the terminal device, and the N preset questions corresponding to the most recent video node are obtained.

[0042] The N preset questions corresponding to the most recent video node are obtained.

[0043] The technical scheme has the beneficial effects that: the most relevant preset questions can be provided for the user according to the playing progress of the explanation video group currently watched by the user. By obtaining the playing progress of the explanation video group being played by the terminal device, the most recent video node can be automatically determined, and N preset questions corresponding to the video node are obtained. In this way, the user can obtain the questions most relevant to the currently watched content, so as to better understand and master the relevant knowledge.

[0044] In some possible implementation manners, the N preset questions corresponding to the most recent video node are obtained by using a question recommendation model.

[0045] The N preset questions corresponding to the most recent video node are obtained by using a question recommendation model.

[0046] The method further includes:

[0047] In response to a click operation on the question button, a text input box and a hold-to-speak button are displayed in the bottom navigation area;

[0048] In response to a text input operation on the text input box, the most recent video node and a personalized question corresponding to the text input operation are stored in a preset message queue in association;

[0049] In response to a long press operation on the hold-to-speak button, speech corresponding to the long press operation is converted into text content, and the most recent video node and a personalized question converted out are stored in the preset message queue in association;

[0050] When the number of messages in the preset message queue is greater than a preset number, all messages in the preset message queue are consumed to update model parameters of a question recommendation model using the messages in the preset message queue, and each message includes a pair of associated video node and personalized question.

[0051] The technical scheme has the beneficial effects that N preset questions corresponding to the most recent video node can be obtained through the question recommendation model, and the model parameters of the question recommendation model can be updated according to the personalized questions of the user. When the user clicks the question button, the text input box and the hold-to-speak button are displayed in the bottom navigation area. If the user feels that the automatically recommended questions cannot meet the needs of the user, the user can propose personalized questions through text input or voice input. These personalized questions are stored in the preset message queue in association with the most recent video node. When the number of messages in the preset message queue is greater than a preset number, all messages in the queue are consumed, and the model parameters of the question recommendation model are updated using the messages. In this way, the question recommendation model can be continuously optimized according to the actual needs of the user, so as to more accurately provide the preset questions corresponding to the explanation video group for the user.

[0052] In some possible implementation manners, the process of updating the model parameters of the question recommendation model using the messages in the preset message queue includes:

[0053] For each message in the preset message queue, the following processing is performed:

[0054] A first sentiment type corresponding to the personalized question in the message is obtained;

[0055] The video node in the message and the first sentiment type are input into the question recommendation model to obtain N predicted questions;

[0056] Based on the personalized question in the message and the N predicted questions, the model parameters of the question recommendation model are updated;

[0057] The utilization problem recommendation model obtains N preset problems corresponding to the last video node, including:

[0058] Obtain a second emotion type corresponding to the user identifier;

[0059] Input the last video node and the second emotion type into the problem recommendation model to obtain N preset problems.

[0060] The beneficial effects of this technical solution are that the model parameters of the problem recommendation model can be updated according to the emotion type of the user, and N preset problems corresponding to the last video node can be obtained according to the emotion type of the user. In the process of updating the problem recommendation model, for each message in the preset message queue, the first emotion type corresponding to the personalized question in the message is obtained, and the video node and the first emotion type are input into the problem recommendation model to obtain N predicted questions. Then, the model parameters of the problem recommendation model are updated based on the prediction error between the personalized question and the predicted question. In the process of obtaining the preset question, the second emotion type corresponding to the user identifier is obtained, and the last video node and the second emotion type are input into the problem recommendation model to obtain N preset questions. In this way, the problem recommendation model can be continuously optimized according to the actual needs and emotional state of the user, thereby better providing problem recommendation services for the user.

[0061] In some possible implementation manners, the updating of the model parameters of the problem recommendation model based on the personalized question in the message and N predicted questions includes:

[0062] The similarity between the personalized question in the message and N predicted questions is calculated respectively;

[0063] The prediction error between the personalized question in the message and the predicted question with the lowest similarity is calculated;

[0064] When the prediction error is greater than a preset error, the model parameters of the problem recommendation model are updated;

[0065] The process of updating the model parameters of the problem recommendation model by using the messages in the preset message queue further includes:

[0066] When the prediction error is not greater than the preset error, the model parameters of the problem recommendation model are not updated.

[0067] The beneficial effects of the technical scheme are that the similarity between the personalized question and the predicted question and the prediction error are calculated to update the model parameters of the question recommendation model. In the updating process, the similarity between the personalized question and the N predicted questions is calculated first, and the prediction error between the personalized question and the predicted question with the lowest similarity is calculated. The prediction error can reflect the accuracy of the question recommendation model in predicting user demand. If the prediction error is large, the accuracy of the question recommendation model is low and needs to be adjusted. If the prediction error is small, the accuracy of the question recommendation model is high and can be continuously used. In this way, the model parameters of the question recommendation model can be updated according to the prediction error, so that the question recommendation model can be continuously optimized according to the actual demand of the user, thereby better recommending the preset question for the user.

[0068] In a second aspect, the present application provides an electronic device for implementing a virtual object interaction function, the electronic device comprising a memory and at least one processor, the memory storing a computer program, and the at least one processor implementing the following steps when executing the computer program:

[0069] In response to an access request of a terminal device, a waiting page is displayed by the terminal device in the process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0070] After the connection is established, a guide page is displayed by the terminal device, the guide page being provided with a guide button;

[0071] In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by the terminal device.

[0072] In a third aspect, the present application provides a computer readable storage medium for implementing a virtual object interaction function, the computer readable storage medium storing a computer program, and the computer program implementing the following steps when executed by at least one processor:

[0073] In response to an access request of a terminal device, a waiting page is displayed by the terminal device in the process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0074] After the connection is established, a guide page is displayed by the terminal device, the guide page being provided with a guide button;

[0075] In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by the terminal device. BRIEF DESCRIPTION OF DRAWINGS

[0076] The application will be further described below with reference to the drawings and embodiments.

[0077] Figure 1 is a flowchart of a virtual object interaction method provided by an embodiment of the application.

[0078] Figure 2 is a flowchart of a method for obtaining a playing page provided by an embodiment of the application.

[0079] Figure 3 is a flowchart of a method for obtaining a plurality of sets of explanation videos provided by an embodiment of the application.

[0080] Figure 4 is a flowchart of a method for automatically recommending a question and explanation provided by an embodiment of the application.

[0081] Figure 5 is a flowchart of a method for obtaining a preset question provided by an embodiment of the application.

[0082] Figure 6 is a flowchart of a method for updating a question recommendation model provided by an embodiment of the application.

[0083] Figure 7 is a flowchart of another method for updating a question recommendation model provided by an embodiment of the application.

[0084] Figure 8 is a flowchart of still another method for updating a question recommendation model provided by an embodiment of the application.

[0085] Figure 9 is a structural block diagram of an electronic device provided by an embodiment of the application.

[0086] Figure 10 shows a structural diagram of a program product provided by an embodiment of the application. DETAILED DESCRIPTION

[0087] The technical solutions in the application will be described below with reference to the drawings and specific embodiments of the application. It should be noted that, under the premise of no conflict, the following described embodiments or technical features can be combined in any manner to form new embodiments.

[0088] In the embodiments of the application, the words such as “exemplary” or “for example” are used to mean serving as an example, instance, or illustration. Any embodiment or design scheme described as “exemplary” or “for example” in the embodiments of the application should not be interpreted as being more preferred or having more advantages than other embodiments or design schemes. Rather, the words such as “exemplary” or “for example” are used in the specific manner to present the relevant concept.

[0089] The technical field and related terms of the embodiments of the present application are briefly described below.

[0090] The virtual object includes a virtual person, a virtual animal, a virtual cartoon image, etc. Among them, the virtual person refers to a personified image constructed by CG technology and running in code form, which has multiple interactive ways such as language communication, expression of expression, and action display. Virtual person technology has rapidly developed in the field of artificial intelligence, and has been applied in many technical fields, such as film and television, media, games, finance, tourism, education, medical treatment, etc. Not only can it customize virtual hosts, virtual anchors, virtual idols, virtual customer service, virtual teachers, virtual lawyers, virtual doctors, etc., but also can generate videos through audio or text content one key.

[0091] Artificial intelligence (AI) is to use digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence researches the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making. Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, automatic driving, intelligent transportation, etc. several major directions.

[0092] Machine learning (ML) is a multi-disciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, etc. Computer programs can learn experience E under a certain category of task T and performance measure P, if its performance in task T can be measured by P, then it will improve with experience E. Machine learning is a branch of computer science that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent, and its applications are widespread in various fields of artificial intelligence.

[0093] Deep learning is a special type of machine learning that learns to represent the world using nested hierarchical concepts, achieving tremendous functionality and flexibility. Each concept is defined as being associated with a simpler one, while more abstract representations are computed in a less abstract manner. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by demonstration.

[0094] Virtual humans (or computer-synthesized characters) refer to the representation of a person's geometric and behavioral characteristics in a computer-generated space (virtual environment). The main function of service-oriented virtual humans is to replace real-life services and provide daily companionship. They are virtualizations of real-world service roles, including virtual anchors, virtual teachers, virtual customer service representatives, and virtual assistants. Their industrial value lies primarily in reducing costs in existing service industries, thus improving efficiency and reducing costs in the existing market. Currently, service-oriented virtual humans (also known as virtual employees) are being implemented in areas such as virtual news anchors, virtual presenters, digital narrators, virtual financial advisors, and virtual lawyers.

[0095] The AI ​​video interaction management backend can configure the virtual human's playback logic, content, and application style as needed, thereby generating an H5 (i.e., HTML5) carrier for interactive virtual human videos that users can see. Because existing service-oriented virtual human interaction processes require playing a series of video clips, a connection needs to be established with a video streaming service. During this connection establishment process, users often experience anxiety and impatience, resulting in a poor interactive experience.

[0096] Based on this, this application provides a virtual object interaction method, an electronic device, and a computer-readable storage medium to improve the prior art.

[0097] The solutions provided in this application involve technologies such as virtual humans, interaction design, artificial intelligence, 3D modeling, and cloud computing, and are specifically illustrated through the following embodiments. It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0098] (Virtual object interaction methods)

[0099] See Figure 1 , Figure 1 This is a flowchart illustrating a virtual object interaction method provided in an embodiment of this application.

[0100] This application provides a virtual object interaction method for implementing virtual object interaction functionality, the method comprising:

[0101] Step S101: in response to an access request of a terminal device, displaying a waiting page by using the terminal device in a process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0102] Step S102: after the connection is established, displaying a guide page by using the terminal device, the guide page being provided with a guide button;

[0103] Step S103: in response to a click operation on the guide button, displaying a play page corresponding to the user identifier by using the terminal device.

[0104] The virtual object interaction method can run on an electronic device, and the electronic device and the terminal device can be independent of each other, or the electronic device and the terminal device can be integrated. When the electronic device and the terminal device are independent of each other, the electronic device can be a computer, a server, or the like.

[0105] In addition, the electronic device and the target server can be independent of each other, or the electronic device and the target server can be integrated.

[0106] The target server can have a function of pushing a stream or pulling a stream, and can push a lecture video group stored in the cloud to the terminal device.

[0107] In the embodiment of the application, the virtual object includes one or more of a virtual person, a virtual animal, and a virtual cartoon image. As an example, the virtual object is a virtual person "JING" (Chinese name: mirror).

[0108] The terminal device is not limited in the embodiment of the application, and can be, for example, a smart terminal device with a display screen, such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a smart wearable device, or the like, or the terminal device can be a workstation or a console with a display screen. The display screen can be a touch display screen or a non-touch display screen.

[0109] The terminal device is not limited in the embodiment of the application, and can be, for example, a smart terminal device with a display screen, such as a mobile phone, a tablet computer, a notebook computer, a desktop computer, a smart wearable device, or the like, or the terminal device can be a workstation or a console with a display screen. The display screen can be a touch display screen or a non-touch display screen.

[0110] When the terminal device requests access, the target server receives the access request and records the user identification, which can be a user account identification or a user device identification, for example. During the connection establishment process, the target server returns a waiting page to the terminal device, prompting the user that the video is being connected. The waiting page can include the name, description, and other information of the virtual object, allowing the user to understand the basic information of the virtual object, shorten the distance between the virtual object and the user, and improve the affinity of the virtual object. Alternatively, the waiting page can include enterprise-related information. After the connection is established, the terminal device displays a guide page. When the user clicks on the guide button on the guide page, the terminal device displays a play page. The guide button can display the words “Explore the Meta Universe”, for example, to guide the user to click.

[0111] The play page is used to play one or more sets of explanation videos corresponding to the virtual object. The user can understand various service information of the enterprise in the explanation video set of the enterprise service virtual person, or learn textbook knowledge and answering skills in the explanation video set of the virtual teacher, or understand medical knowledge such as disease symptoms in the explanation video set of the virtual doctor, or understand the historical background, cultural significance, and artistic value of the museum's collection in the explanation video set of the virtual museum guide.

[0112] Thus, in the connection establishment process, the waiting page is used to alleviate the anxiety of the user during the waiting period, improving the interactive experience. After the connection is established, the guide page and the guide button are set to clearly and intuitively guide the user to enter the play page, improving the smoothness and comfort of user access. On the other hand, the connection between the terminal device and the target server is established, allowing the user to more smoothly watch the explanation video set and more conveniently select the explanation video set, improving the user experience. On the other hand, the corresponding play page is displayed according to the user identification, allowing the play page to be personalized and better improving the user's participation and stickiness, thereby enhancing the user's interest and attention to the explanation video set and improving the user's stickiness to the virtual object interactive mode. On the other hand, a separate guide page and guide button are set, and the user will enter the play page of the explanation video set only after a click operation. When the user triggers an access request by mistake, the explanation video set will not automatically play, avoiding unnecessary aversion. That is, a threshold is set for entering the play page, improving the ritual sense of the virtual object interactive process and maintaining the commercial value of the virtual object.

[0113] In some embodiments, the waiting page is provided with a connection progress bar and a carousel area;

[0114] The connection progress bar is used to display the connection progress in real time;

[0115] The carousel area is used to carousel a plurality of carousel images, or the carousel area is used to carousel a plurality of carousel videos.

[0116] As an example, the connection progress bar can be set below or in the middle of the waiting page, and the carousel area can be set in the middle or above the waiting page accordingly. The connection progress can be represented in percentage form, located to the right of the connection progress bar.

[0117] In this way, the connection progress bar and the carousel area are set in the waiting page, allowing users to understand the connection progress in real time during the process of establishing a connection, and attracting the user's attention by rotating multiple carousel images or carousel videos, improving the user experience. On the one hand, the waiting page sets the connection progress bar, allowing users to understand the progress of establishing a connection in real time, thereby reducing the user's anxiety while waiting; and the carousel area can display multiple images or videos to provide users with more information and entertainment, thereby improving user satisfaction and experience. On the other hand, by designing a sophisticated waiting page that showcases the professionalism and innovation of a company or product, the brand image and brand awareness can be enhanced, and the company or product can gain more praise and word-of-mouth. On the other hand, displaying attractive images or videos in the carousel area can attract the user's attention, increasing the user's attention and interaction, thereby improving marketing effectiveness and conversion rates. On the other hand, by designing a modular waiting page, different parts of the page can be easily maintained and updated, using a modular design, high configurability, unified management, easy extension and update, improving the maintainability and scalability of the waiting page, reducing maintenance costs and improving development efficiency. On the other hand, according to different application scenarios, the style and content of the connection progress bar and the carousel area can be customized to meet different user needs and application scenarios, improving the flexibility and adaptability of the application.

[0118] The waiting page can be divided into multiple modules, such as the connection progress bar and the carousel area, each of which can be independently designed, developed and maintained, thereby reducing maintenance costs and improving development efficiency. The design and content of the waiting page can be configured and managed through configuration files, thereby facilitating subsequent updates and maintenance. For example, the style and color of the connection progress bar, the content of the carousel area such as images or videos, etc. can be set through configuration files. By designing reusable components, the design and development of the waiting page can be managed uniformly, thereby improving the consistency and stability of the page, reducing maintenance costs and improving development efficiency. The waiting page can be extended and customized as needed, such as adding new modules, modifying styles and layouts, etc., thereby adapting to different needs and application scenarios. By decoupling the design and content of the waiting page from the backend data, updates and maintenance can be easily performed, such as modifying the progress of the connection progress bar, updating the images or videos in the carousel area, etc. At the same time, updating the content can also be achieved through configuration files, API interfaces, etc., thereby simplifying the update process.

[0119] Referring to Figure 2 , Figure 2 is a flowchart of a process of obtaining a playing page provided by an embodiment of the present application.

[0120] In some embodiments, the process of obtaining the playing page corresponding to the user identifier comprises:

[0121] Step S201: Obtain the playing type corresponding to the user identifier.

[0122] Step S202: When the playing type is not played, take an initial playing page as the playing page corresponding to the user identifier, the initial playing page being used to play a plurality of explanation video groups in a preset order.

[0123] Step S203: When the playing type is playing interruption, take an interruption playing page of the last time of interruption playing as the playing page corresponding to the user identifier, the interruption playing page being used to play the explanation video group corresponding to the interruption playing page and the explanation video groups after it in the preset order.

[0124] Step S204: When the playing type is playing end, take an end playing page as the playing page corresponding to the user identifier.

[0125] Thus, different playing pages can be provided for the user according to the playing type (i.e. the playing progress type) of the user. When the playing type of the user is not played, the terminal device displays the initial playing page to let the user play the plurality of explanation video groups in the preset order. When the playing type of the user is playing interruption, the terminal device displays the interruption playing page of the last time of interruption playing to let the user continue to play the explanation video group corresponding to the interruption playing page and the explanation video groups after it in the preset order. When the playing type of the user is playing end, the terminal device displays the end playing page. In this way, the appropriate playing page can be reached according to the playing progress of the user, so that a better watching experience can be obtained. By recording the playing state and position of the user, playing the plurality of explanation video groups in the preset order, and supporting the function of watching the subsequent part after interruption playing, the watching experience of the user can be improved, the video watching time length can be increased, and the video playing efficiency can be improved, so that the use value and competitiveness of the platform can be improved.

[0126] The embodiment of the application can determine the playing state of the user according to the user identifier, to ensure that the user can continue to watch the video at the last interrupted position or watch the next video from the last watched position. This playing mode can improve the user's watching experience and avoid the need for the user to manually find the last interrupted position. By playing the multiple explanation video groups in a preset order and supporting functions such as subsequent watching after interruption, the user can more conveniently watch multiple videos, thereby increasing the video watching time and improving the use value of the platform. By recording the last played position of the user, the user need not manually find the last watched position, thereby improving the video playing efficiency. Meanwhile, the multiple explanation video groups are played in a preset order, so that the user need not manually select the next video, saving the user operation time and improving the use efficiency of the platform.

[0127] In some other embodiments, when the playing type is playing interruption, the user can not be directly led to the interruption playing page, but be asked whether to continue playing from the explanation video group corresponding to the interruption playing page, and then based on the user's selection, it is determined whether to continue playing from the explanation video group corresponding to the interruption playing page or to play the multiple explanation video groups from the beginning.

[0128] Referring to Figure 3 , Figure 3 is a flowchart for obtaining multiple explanation video groups provided by the embodiment of the application.

[0129] In some embodiments, the process of obtaining the multiple explanation video groups includes:

[0130] Step S301: Obtain the device configuration information and the network configuration information corresponding to the terminal device;

[0131] Step S302: Obtain the definition type corresponding to the device configuration information and the network configuration information;

[0132] Step S303: Obtain the multiple explanation video groups corresponding to the definition type.

[0133] The device configuration information may, for example, include one or more of brand and model, operating system version, processor type and frequency, memory and storage capacity, screen resolution and size, communication technology and network standard.

[0134] The network configuration information may, for example, include network bandwidth, network delay, network connection quality, etc.

[0135] The definition type may, for example, include smooth (corresponding to a code rate of 0.5M), standard definition (corresponding to a code rate of 1M), high definition (corresponding to a code rate of 2M), super definition, etc. Alternatively, the definition type may, for example, include 480P, 720P, 1080P, 4K, 8K, etc.

[0136] In some embodiments, the obtaining the corresponding definition type of the device configuration information and the network configuration information (i.e., step S302) comprises inputting the device configuration information and the network configuration information into a definition model to obtain the corresponding definition type. The definition model can be obtained by training a preset deep learning model using a training set.

[0137] Thus, the appropriate definition type can be automatically selected according to the configuration information of the terminal device and the network configuration information, thereby obtaining multiple sets of explanation videos corresponding to the definition type. This can improve the user experience and avoid the problem of unsmooth video playback due to poor device configuration or network conditions. On the one hand, the appropriate definition type is automatically selected according to the device configuration information and the network configuration information of the user, thereby providing a better viewing experience. On the other hand, by selecting the appropriate definition type, the problems of video lag, slow loading, etc. can be avoided, thereby improving the viewing quality. On the other hand, by selecting the appropriate definition type, the data transmission amount can be reduced, thereby reducing the user's traffic consumption. On the other hand, by providing personalized video viewing experience, the user's satisfaction and trust in the product or service can be improved, thereby enhancing the user's stickiness. On the other hand, by automatically selecting the definition type, manual intervention and adjustment can be reduced, thereby reducing the maintenance cost of the product or service.

[0138] In some embodiments, each of the playback pages is provided with a definition navigation area.

[0139] As an example, the definition navigation area is arranged at the upper left of the playback page, and the definition navigation area is provided with a definition adjustment button. By clicking the definition adjustment button, the user can select and switch the definition.

[0140] As an example, the user can be automatically prompted to switch the definition according to the real-time network status (such as packet loss condition, etc.).

[0141] In some embodiments, each of the playback pages is provided with a side navigation area and / or a bottom navigation area.

[0142] The side navigation area is provided with one or more of the following function buttons: table of contents, subscription, sharing, form, trial, cooperation, and external link jump.

[0143] The bottom navigation area is provided with one or more of the following function buttons: pause / resume, table of contents, mode switching, speak and ask.

[0144] Thus, the side navigation area and / or the bottom navigation area are set for each playing page, and rich function buttons are provided to enable the user to have a better viewing experience. The function buttons of the side navigation area include directory, subscription, sharing, form, trial, cooperation, and external link jump, etc., which can meet various needs of the user when watching the explanation video group. For example, the user can quickly jump to the video node that the user wants to watch by clicking the directory button, subscribe to the channel of interest by clicking the subscription button, or guide the user to subscribe to the public number, the up master, the blogger, etc., share the video with friends by clicking the sharing button, fill in the related form by clicking the form button, try the related product by clicking the trial button, cooperate with the video provider by clicking the cooperation button, and jump to the related webpage by clicking the external link jump button. In addition, the bottom navigation area provides function buttons such as pause / resume, directory, mode switching, hold to talk, and ask questions, etc., to provide more control and interaction modes for the user, which facilitates the user to operate when watching the video, and the user can more flexibly control the video playing process, and can ask questions at any time.

[0145] In some embodiments, the mode switching can realize switching between the explanation mode and the question and answer mode. In the explanation mode, the virtual person is dominant, and the corresponding explanation video group is played based on the video node selected by the user. In the question and answer mode, the user is dominant, and the corresponding explanation video group is automatically matched and played based on the question (including the preset question and the personalized question) asked by the user.

[0146] When the user clicks the hold to talk button, the audio acquisition component of the terminal device can be called to realize the audio acquisition function to obtain the voice input by the user. In actual application, the user can long press the hold to talk button to realize continuous acquisition of the voice. The single voice input time length can be set to an upper limit time length, for example, 10 seconds, 30 seconds, 1 minute, or 3 minutes, etc.

[0147] Referring to Figure 4 , Figure 4 is a flowchart of the process of automatically recommending questions and explanation provided by the embodiments of the present application.

[0148] In some embodiments, the bottom navigation area is provided with a question button;

[0149] The method further includes:

[0150] Step S401: In response to a click operation on the question button, an explanation video group being played by the terminal device is acquired;

[0151] Step S402: Based on the explanation video group being played by the terminal device, N preset questions corresponding to the explanation video group are acquired, and N is a positive integer;

[0152] Step S403: displaying N preset questions in the form of a floating layer by using the terminal device;

[0153] Step S404: in response to a click operation on one of the preset questions, obtaining an explanation video set corresponding to the clicked preset question;

[0154] Step S405: playing the explanation video set corresponding to the clicked preset question by using the terminal device.

[0155] The present embodiment does not limit N, which may be, for example, 1, 2, 3, 5, 10, 15, 20, etc.

[0156] Thus, by clicking the question button in the bottom navigation area, the user can obtain N preset questions corresponding to the explanation video set being played by the terminal device, and the preset questions are displayed in the form of a floating layer. The user can click one of the questions to obtain the corresponding explanation video set. In this way, the user can quickly find the information they need, thereby better understanding and mastering the relevant knowledge. On the one hand, the user can quickly obtain the required information and knowledge by clicking the question button, and can quickly understand the key points of the video content through the preset questions, thereby improving the learning efficiency. On the other hand, by providing preset questions and corresponding explanation video sets, the user can more vividly and intuitively understand the video content, thereby improving the interest of learning and improving the user's interest in learning. On the other hand, the user is allowed to click the question button and obtain the corresponding preset questions according to the played explanation video set, thereby improving the user's participation, and the user can more actively participate in learning and exploration. On the other hand, by displaying the preset questions in the form of a floating layer in the terminal device, the user can more conveniently obtain the required information without leaving the current playing page, thereby improving the user experience.

[0157] Referring to Figure 5 , Figure 5 is a flowchart of obtaining preset questions provided by the present embodiment.

[0158] In some embodiments, each explanation video set is provided with one or more video nodes;

[0159] The obtaining of N preset questions corresponding to the explanation video set being played by the terminal device (i.e., step S105) comprises:

[0160] Step S501: obtaining a latest video node based on the playing progress of the explanation video set being played by the terminal device;

[0161] Step S502: obtaining N preset questions corresponding to the latest video node.

[0162] Each explanation video group includes one or more videos connected in series, and each explanation video group can be provided with one or more video nodes, with the starting position of each video as a video node.

[0163] Thus, the most relevant preset question can be provided for the user according to the playing progress of the explanation video group currently watched by the user. By obtaining the playing progress of the explanation video group being played by the terminal device, the most recent video node can be determined automatically, and the N preset questions corresponding to the video node can be obtained. In this way, the user can obtain the question most relevant to the currently watched content, thereby better understanding and mastering the relevant knowledge.

[0164] The most recent video node refers to the most recent video node found from the current playing progress, that is, the most recent video node is found from the current playing progress (past) rather than from the future.

[0165] As an example, the terminal device is playing the explanation video group 03#, which has a duration of 10 minutes and is provided with five video nodes at 0 minutes, 2 minutes, 4 minutes, 6 minutes and 8 minutes. When the playing progress is 5 minutes, the most recent video node refers to the video node at 4 minutes (rather than the video node at 6 minutes); when the playing progress is 7 minutes and 30 seconds, the most recent video node refers to the video node at 6 minutes (rather than the video node at 8 minutes).

[0166] Referring to Figure 6 , Figure 6 is a flowchart of a process for updating a question recommendation model provided by an embodiment of the present application.

[0167] In some embodiments, the obtaining of the N preset questions corresponding to the most recent video node (i.e., step S502) includes:

[0168] obtaining the N preset questions corresponding to the most recent video node by using the question recommendation model;

[0169] In addition to steps S101-S103, the method further includes:

[0170] Step S601: in response to a click operation on the question button, displaying a text input box and a hold-to-speak button in the bottom navigation area;

[0171] Step S602: in response to a text input operation on the text input box, storing the most recent video node and the personalized question corresponding to the text input operation in a preset message queue;

[0172] Step S603: In response to the long press operation on the hold-to-speak button, converting the voice corresponding to the long press operation into text content, and storing the most recent video node and the converted personalized question in association into the preset message queue.

[0173] Step S604: When the number of messages in the preset message queue is greater than a preset number, consuming all the messages in the preset message queue to update the model parameters of the question recommendation model by using the messages in the preset message queue, each of the messages including a pair of associated video node and personalized question.

[0174] The preset number is not limited in the embodiments of the present application, and may be, for example, 10, 100, 1000, etc.

[0175] Thus, the N preset questions corresponding to the most recent video node can be obtained by the question recommendation model, and the model parameters of the question recommendation model can be updated according to the personalized questions of the user. When the user clicks the question button, the text input box and the hold-to-speak button are displayed in the bottom navigation area, and if the user feels that the automatically recommended questions cannot meet the needs of the user, the user can propose personalized questions through text input or voice input. These personalized questions are stored in association with the most recent video node into the preset message queue. When the number of messages in the preset message queue is greater than a preset number, the system consumes all the messages in the queue, and updates the model parameters of the question recommendation model by using the messages. In this way, the question recommendation model can be continuously optimized according to the actual needs of the user, so as to more accurately provide the preset questions corresponding to the lecture video group.

[0176] In some other embodiments, in addition to steps S101-S103, the method further includes:

[0177] In response to the click operation on the question button, displaying a text input box in the bottom navigation area;

[0178] In response to the text input operation on the text input box, storing the most recent video node and the personalized question corresponding to the text input operation in association into a preset message queue;

[0179] When the number of messages in the preset message queue is greater than a preset number, consuming all the messages in the preset message queue to update the model parameters of the question recommendation model by using the messages in the preset message queue, each of the messages including a pair of associated video node and personalized question.

[0180] In some other embodiments, in addition to steps S101-S103, the method further includes:

[0181] In response to a click operation on the question button, a hold-to-speak button is displayed in the bottom navigation area.

[0182] In response to a long press operation on the hold-to-speak button, the voice corresponding to the long press operation is converted into text content, and the most recent video node and the converted personalized question are associated and stored in the preset message queue.

[0183] When the number of messages in the preset message queue is greater than a preset number, all messages in the preset message queue are consumed to update the model parameters of the question recommendation model using the messages in the preset message queue. Each message includes an associated video node and a personalized question.

[0184] Referring to Figure 7 , Figure 7 is another flowchart for updating a question recommendation model provided by an embodiment of the present application.

[0185] In some embodiments, the process of updating the model parameters of the question recommendation model using the messages in the preset message queue includes:

[0186] For each message in the preset message queue, the following processing is performed:

[0187] Step S701: Obtain a first sentiment type corresponding to the personalized question in the message;

[0188] Step S702: Input the video node in the message and the first sentiment type into the question recommendation model to obtain N predicted questions;

[0189] Step S703: Update the model parameters of the question recommendation model based on the personalized question in the message and the N predicted questions;

[0190] The N preset questions corresponding to the most recent video node obtained using the question recommendation model include:

[0191] Obtain a second sentiment type corresponding to the user identifier;

[0192] Input the most recent video node and the second sentiment type into the question recommendation model to obtain N preset questions.

[0193] In some embodiments, the acquiring the second emotion type corresponding to the user identifier comprises: acquiring a user image by using an image acquisition component of the terminal device; and identifying an emotion type corresponding to the user image as the second emotion type corresponding to the user identifier. The emotion type corresponding to the user image can be identified, for example, by inputting the user image into a trained emotion classification model to identify the corresponding emotion type.

[0194] In other embodiments, the acquiring the second emotion type corresponding to the user identifier comprises: acquiring, by using the terminal device, an emotion type input or selected by the user as the second emotion type corresponding to the user identifier.

[0195] In some embodiments, the emotion type can be divided into positive, negative, and neutral, for example. The positive emotion can be further divided into like, appreciate, happy, satisfied, grateful, excited, joyful, cheerful, and satisfied, for example. The negative emotion can be further divided into annoyed, lost, painful, depressed, angry, sad, anxious, worried, and desperate, for example. The neutral emotion can be further divided into calm, objective, indifferent, flat, indifferent, and neutral, for example.

[0196] In other embodiments, the emotion type can be divided into positive, negative, and neutral, for example, without limitation in the present application.

[0197] It should be noted that the first emotion type and the second emotion type herein are both emotion types, and the prefixes of “first” and “second” serve to distinguish.

[0198] Therefore, the model parameters of the question recommendation model can be updated according to the emotion type of the user, and the N preset questions corresponding to the most recent video node can be acquired according to the emotion type of the user. In the process of updating the question recommendation model, for each message in the preset message queue, the first emotion type corresponding to the personalized question in the message is acquired, and the video node and the first emotion type are input into the question recommendation model to obtain N predicted questions. Then, the model parameters of the question recommendation model are updated based on the prediction error between the personalized question and the predicted question. In the process of acquiring the preset question, the second emotion type corresponding to the user identifier is acquired, and the most recent video node and the second emotion type are input into the question recommendation model to obtain N preset questions. In this way, the question recommendation model can be continuously optimized according to the actual needs and emotional state of the user, thereby better providing question recommendation services for the user.

[0199] Referring to Figure 8 , Figure 8 is another flowchart of updating the question recommendation model provided by the embodiments of the present application.

[0200] In some embodiments, the updating (i.e., step S703) of the model parameters of the question recommendation model based on the personalized question in the message and the N predicted questions comprises:

[0201] Step S801: calculating the similarity between the personalized question in the message and the N predicted questions, respectively;

[0202] Step S802: calculating the prediction error between the personalized question in the message and the predicted question with the lowest similarity;

[0203] Step S803: updating the model parameters of the question recommendation model when the prediction error is greater than the preset error;

[0204] The process of updating the model parameters of the question recommendation model using the messages in the preset message queue further comprises:

[0205] When the prediction error is not greater than the preset error, the model parameters of the question recommendation model are not updated.

[0206] Thus, the model parameters of the question recommendation model are updated by calculating the similarity and the prediction error between the personalized question and the predicted question. In the updating process, the similarity between the personalized question and the N predicted questions is calculated first, and the prediction error between the personalized question and the predicted question with the lowest similarity is calculated. The prediction error can reflect the accuracy of the question recommendation model in predicting user demand. If the prediction error is large, it means that the accuracy of the question recommendation model is low and needs to be adjusted. If the prediction error is small, it means that the accuracy of the question recommendation model is high and can be used continuously. In this way, the model parameters of the question recommendation model can be updated according to the prediction error, so that the question recommendation model can be continuously optimized according to the actual demand of the user, thereby better recommending the preset question to the user.

[0207] The embodiments of the present application can more accurately understand the demand and emotional tendency of the user by classifying the personalized question in the message, thereby improving the accuracy and reliability of question recommendation. By inputting the video node and the first emotional type in the message into the question recommendation model, a group of predicted questions consistent with the demand and emotion of the user can be obtained, thereby increasing the diversity and flexibility of question recommendation. By calculating the similarity between the personalized question and the predicted question, the recommended questions can be further filtered and sorted to ensure the matching degree of the recommended questions and the demand of the user. By calculating the prediction error between the personalized question and the predicted question with the lowest similarity, and updating the model parameters of the question recommendation model according to whether the error value is greater than the preset error, the accuracy and stability of the model can be improved, thereby further optimizing the effect of question recommendation.

[0208] In one specific application scenario, the embodiment of the present application further provides a virtual object interaction method for realizing a virtual object interaction function, the method comprising:

[0209] In response to an access request of a terminal device, a waiting page is displayed by using the terminal device in the process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0210] After the connection is established, a guide page is displayed by using the terminal device, the guide page being provided with a guide button;

[0211] In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by using the terminal device.

[0212] The waiting page is provided with a connection progress bar and a carousel area; the connection progress bar is used to display a connection progress in real time; and the carousel area is used to carousel a plurality of carousel images, or the carousel area is used to carousel a plurality of carousel videos.

[0213] Each of the play pages is provided with a side navigation area and / or a bottom navigation area; the side navigation area is provided with one or more of the following function buttons: catalog, subscription, sharing, form, trial, cooperation and external link jump; and the bottom navigation area is provided with one or more of the following function buttons: pause / resume, catalog, mode switching, hold to speak and ask questions.

[0214] The process of obtaining the play page corresponding to the user identifier comprises:

[0215] A play type corresponding to the user identifier is obtained.

[0216] When the play type is not played, an initial play page is taken as the play page corresponding to the user identifier, the initial play page being used to play a plurality of explanation video groups in a preset order;

[0217] When the play type is interrupted, an interrupted play page of the last time of interrupted play is taken as the play page corresponding to the user identifier, the interrupted play page being used to play an explanation video group corresponding to the interrupted play page and an explanation video group after the explanation video group in the preset order;

[0218] When the play type is played, an end play page is taken as the play page corresponding to the user identifier.

[0219] The process of obtaining a plurality of the explanation video groups comprises:

[0220] Device configuration information and network configuration information corresponding to the terminal device are obtained.

[0221] obtaining a definition type corresponding to the device configuration information and the network configuration information;

[0222] obtaining a plurality of the explanation video groups corresponding to the definition type.

[0223] The method further comprises the following automatic recommendation of questions and explanation processes:

[0224] In response to a click operation on the question button, obtaining an explanation video group being played by the terminal device;

[0225] Based on the playing progress of the explanation video group being played by the terminal device, obtaining a latest video node;

[0226] obtaining a second emotion type corresponding to the user identifier;

[0227] inputting the latest video node and the second emotion type into the question recommendation model to obtain N preset questions, N being a positive integer; each of the explanation video groups is provided with one or more video nodes;

[0228] displaying the N preset questions in the form of a floating layer by using the terminal device;

[0229] In response to a click operation on one of the preset questions, obtaining an explanation video group corresponding to the clicked preset question;

[0230] playing the explanation video group corresponding to the clicked preset question by using the terminal device.

[0231] The method further comprises the following model updating process:

[0232] In response to a click operation on the question button, displaying a text input box and a hold-to-speak button in the bottom navigation area;

[0233] In response to a text input operation on the text input box, storing an individualized question corresponding to the text input operation and the latest video node in a preset message queue in association;

[0234] In response to a long press operation on the hold-to-speak button, converting a voice corresponding to the long press operation into text content, and storing the individualized question converted and the latest video node in the preset message queue in association;

[0235] When the number of messages in the preset message queue is greater than a preset number, all messages in the preset message queue are consumed to update model parameters of the question recommendation model by using the messages in the preset message queue, each of the messages comprising a pair of associated video node and personalized question.

[0236] The process of updating the model parameters of the question recommendation model by using the messages in the preset message queue comprises:

[0237] For each of the messages in the preset message queue, the following processing is performed:

[0238] A first sentiment type corresponding to the personalized question in the message is obtained;

[0239] The video node in the message and the first sentiment type are input into the question recommendation model to obtain N predicted questions;

[0240] Similarities between the personalized question in the message and the N predicted questions are respectively calculated;

[0241] A prediction error between the personalized question in the message and the predicted question with the lowest similarity is calculated;

[0242] When the prediction error is greater than a preset error, the model parameters of the question recommendation model are updated;

[0243] When the prediction error is not greater than the preset error, the model parameters of the question recommendation model are not updated.

[0244] (electronic device)

[0245] Embodiments of the present application also provide an electronic device, the specific embodiments of which are consistent with the embodiments described in the above method embodiments and the technical effects achieved, and some contents will not be described again.

[0246] The electronic device is used to implement a virtual object interaction function, and comprises a memory and at least one processor, the memory stores a computer program, and the at least one processor implements the following steps when executing the computer program:

[0247] In response to an access request of a terminal device, a waiting page is displayed by using the terminal device in a process of establishing a connection between the terminal device and a target server, the access request being used to indicate a user identifier;

[0248] After the connection is established, a guide page is displayed by using the terminal device, and the guide page is provided with a guide button;

[0249] In response to a click operation on the guide button, the terminal device displays a play page corresponding to the user identifier.

[0250] In some embodiments, the waiting page is provided with a connection progress bar and a carousel area.

[0251] The connection progress bar is used to display the connection progress in real time.

[0252] The carousel area is used to carousel a plurality of carousel images, or the carousel area is used to carousel a plurality of carousel videos.

[0253] In some embodiments, the at least one processor executes the computer program in the following manner to obtain the play page corresponding to the user identifier:

[0254] Obtain the play type corresponding to the user identifier.

[0255] When the play type is not played, an initial play page is used as the play page corresponding to the user identifier, and the initial play page is used to play a plurality of explanation video groups in a preset order.

[0256] When the play type is interrupted, an interrupted play page corresponding to the last time of interrupted play is used as the play page corresponding to the user identifier, and the interrupted play page is used to play the explanation video group corresponding to the interrupted play page and the explanation video group after it in the preset order.

[0257] When the play type is played, an end play page is used as the play page corresponding to the user identifier.

[0258] In some embodiments, the at least one processor executes the computer program in the following manner to obtain a plurality of explanation video groups:

[0259] Obtain the device configuration information and the network configuration information corresponding to the terminal device.

[0260] Obtain the definition type corresponding to the device configuration information and the network configuration information.

[0261] Obtain a plurality of explanation video groups corresponding to the definition type.

[0262] In some embodiments, each of the play pages is provided with a side navigation area and / or a bottom navigation area.

[0263] The side navigation area is provided with one or more of the following function buttons: directory, subscription, sharing, form, trial, cooperation, and external link jump.

[0264] The bottom navigation area is provided with one or more of the following function buttons: pause / resume, catalog, mode switching, hold talking and asking questions.

[0265] In some embodiments, the bottom navigation area is provided with a question asking button;

[0266] The at least one processor, when executing the computer program, further implements the following steps:

[0267] In response to a click operation on the question asking button, a set of explanation videos being played by the terminal device is obtained;

[0268] Based on the set of explanation videos being played by the terminal device, N preset questions corresponding thereto are obtained, N being a positive integer;

[0269] The terminal device displays the N preset questions in the form of a floating layer;

[0270] In response to a click operation on one of the preset questions, a set of explanation videos corresponding to the clicked preset question is obtained;

[0271] The terminal device plays the set of explanation videos corresponding to the clicked preset question.

[0272] In some embodiments, each set of explanation videos is provided with one or more video nodes;

[0273] The at least one processor, when executing the computer program, obtains N preset questions corresponding to the set of explanation videos being played by the terminal device in the following manner:

[0274] Based on the playing progress of the set of explanation videos being played by the terminal device, a latest video node is obtained;

[0275] N preset questions corresponding to the latest video node are obtained.

[0276] In some embodiments, the at least one processor, when executing the computer program, obtains N preset questions corresponding to the latest video node in the following manner:

[0277] A question recommendation model is used to obtain N preset questions corresponding to the latest video node;

[0278] The at least one processor, when executing the computer program, further implements the following steps:

[0279] In response to a click operation on the question asking button, a text input box and a hold talking button are displayed in the bottom navigation area;

[0280] In response to a text input operation on the text input box, the most recent video node and a personalized question corresponding to the text input operation are stored in a preset message queue in association;

[0281] In response to a long press operation on the hold-to-speak button, speech corresponding to the long press operation is converted into text content, and the most recent video node and a personalized question converted from the speech are stored in the preset message queue in association;

[0282] When the number of messages in the preset message queue is greater than a preset number, all messages in the preset message queue are consumed to update model parameters of a question recommendation model by using the messages in the preset message queue, and each message includes a pair of associated video node and personalized question.

[0283] In some embodiments, the at least one processor updates the model parameters of the question recommendation model by using the messages in the preset message queue when executing the computer program in the following manner:

[0284] For each message in the preset message queue, the following processing is performed:

[0285] A first sentiment type corresponding to the personalized question in the message is obtained;

[0286] The video node in the message and the first sentiment type are input into the question recommendation model to obtain N predicted questions;

[0287] The model parameters of the question recommendation model are updated based on the personalized question in the message and the N predicted questions;

[0288] The N preset questions corresponding to the most recent video node are obtained by using the question recommendation model, including:

[0289] A second sentiment type corresponding to the user identifier is obtained;

[0290] The most recent video node and the second sentiment type are input into the question recommendation model to obtain the N preset questions.

[0291] In some embodiments, the at least one processor updates the model parameters of the question recommendation model based on the personalized question in the message and the N predicted questions when executing the computer program in the following manner:

[0292] Similarities between the personalized question in the message and the N predicted questions are calculated respectively;

[0293] calculate a prediction error between the personalized question in the message and the predicted question with the lowest similarity;

[0294] update the model parameters of the question recommendation model when the prediction error is greater than a preset error;

[0295] The at least one processor executes the computer program and further updates the model parameters of the question recommendation model by using the message in the preset message queue in the following manner:

[0296] When the prediction error is not greater than the preset error, the model parameters of the question recommendation model are not updated.

[0297] Referring to Figure 9 , Figure 9 is a structural block diagram of an electronic device 10 provided by an embodiment of the present application.

[0298] The electronic device 10 may, for example, include at least one memory 11, at least one processor 12, and a bus 13 connecting different platform systems.

[0299] The memory 11 can include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 111 and / or a cache memory 112, and can further include a read-only memory (ROM) 113.

[0300] The memory 11 further stores a computer program, which can be executed by the processor 12 to enable the processor 12 to implement the steps of any of the above methods.

[0301] The memory 11 can further include a utility 114 having at least one program module 115, such as an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples can include implementation of a network environment.

[0302] Correspondingly, the processor 12 can execute the above computer program and can execute the utility 114.

[0303] The processor 12 can employ one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.

[0304] The bus 13 can be one or more of several types of bus structures including a memory bus or memory controller, a peripheral bus, a graphics acceleration bus, a processor or local bus using any of a variety of bus architectures, or a combination thereof.

[0305] The electronic device 10 can also communicate with one or more external devices such as a keyboard or pointing device, a Bluetooth device, etc.; other devices associated with the electronic device 10; and / or one or more devices that enable provide connectivity to the electronic device 10 to one or more other computing devices; and / or any devices (e.g., a router, a modem, etc.) that enable the electronic device 10 to communicate in a networked environment. Such communication can be facilitated by an Input / Output Interface 14. Still yet, the electronic device 10 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or the Internet, through a network adapter 15. The network adapter 15 can be communicatively coupled to the other components of the electronic device 10 through the bus 13. It should be appreciated that the network adapter 15 and / or the bus 13 can be implemented using one or more types of technology, including, but not limited to, Ethernet, Token Ring, FDDI, Wi-Fi, IEEE 1394, Bluetooth, etc.

[0306] (computer-readable storage medium)

[0307] The embodiment of the present application further provides a computer readable storage medium, the computer readable storage medium stores a computer program, the computer program is executed by at least one processor to realize the steps of any one of the above methods or realize the functions of any one of the above electronic devices, and the specific embodiments and the achieved technical effects are the same as those described in the above method embodiments, and part of the content will not be described herein.

[0308] Referring to Figure 10 , Figure 10 A structure schematic diagram of a program product provided by the embodiment of the present application is shown.

[0309] The program product is used to implement any of the above methods. The program product can take a portable compact disc read-only memory (CD-ROM) and include program code, and can be run on a terminal device, such as a personal computer. However, the program product of the present application is not limited to this, and in the embodiments of the present application, the readable storage medium can be any tangible medium containing or storing a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. The program product can take any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0310] The computer readable storage medium can include a data signal transported over a carrier wave and can be baseband or propagated along with carriers. The propagated carrier can take any of a variety of forms, including, but not limited to electro-magnetic, optical, or any suitable combination thereof. A computer readable storage medium can be any medium that can be read by an instruction execution system, apparatus, or device, or that can communicate a program to such a system, apparatus, or device for use by or in connection with the instruction execution system, apparatus, or device. The program code contained on the computer readable storage medium can be transmitted using any suitable medium, including, but not limited to wireless, wire line, optical fiber cable, RF, etc., or any suitable combination of the above. Program code embodied on a computer readable storage medium can be transmitted using any combination of one or more program languages, including, but not limited to object oriented program languages, such as Java, C++, etc., as well as conventional procedural program languages, such as the C language, or similar program languages. The program code can execute entirely on a user's computing device, partly on a user's computing device, as a stand-alone software package, partly on a user's computing device and partly on a remote computing device, or entirely on a remote computing device or server. In the latter scenario, the remote computing device can be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computing device, such as through the Internet using an Internet service provider.

[0311] The application is described from the use purpose, efficiency, progress and novelty, which meets the function improvement and use requirements emphasized by the patent law. The above description and drawings are only the preferred embodiments of the application, and are not limited to the application. Therefore, all similar, similar, equivalent replacement or modification, etc. made within the scope of the patent application of the application, shall be within the scope of the patent application protection of the application.

Claims

1. A virtual object interaction method, characterized by, The method for realizing the virtual object interaction function comprises the following steps: In response to an access request of a terminal device, a waiting page is displayed by using the terminal device in the process of establishing a connection between the terminal device and a target server, and the access request is used to indicate a user identifier; After the connection is established, a guide page is displayed by using the terminal device, and the guide page is provided with a guide button; In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by using the terminal device; Each play page is provided with a question button; In response to a click operation on the question button, a set of explanation videos being played by the terminal device is acquired; Based on the set of explanation videos being played by the terminal device, N preset questions corresponding to the set of explanation videos are acquired, and N is a positive integer; In response to a click operation on one of the preset questions, a set of explanation videos corresponding to the clicked preset question is acquired, and each set of explanation videos is provided with one or more video nodes; The set of explanation videos corresponding to the clicked preset question is played by using the terminal device; The step of acquiring, based on the set of explanation videos being played by the terminal device, N preset questions corresponding to the set of explanation videos comprises the following step: acquiring a latest video node based on a playing progress of the set of explanation videos being played by the terminal device; N preset questions corresponding to the latest video node are acquired; The step of acquiring N preset questions corresponding to the latest video node comprises the following step: acquiring, by using a question recommendation model, N preset questions corresponding to the latest video node; The step of acquiring, by using the question recommendation model, N preset questions corresponding to the latest video node comprises the following steps: acquiring a second emotional type corresponding to the user identifier; The latest video node and the second emotional type are input into the question recommendation model to obtain N preset questions. The step of acquiring the second emotional type corresponding to the user identifier comprises the following steps: A user image is acquired by using an image acquisition component of the terminal device, and an emotional type corresponding to the user image is identified as the second emotional type corresponding to the user identifier.

2. The virtual object interaction method of claim 1, wherein, The waiting page is provided with a connection progress bar and a carousel area; The connection progress bar is used to display a connection progress in real time; The carousel area is used to carousel a plurality of carousel images, or the carousel area is used to carousel a plurality of carousel videos.

3. The virtual object interaction method of claim 1, wherein, The process of acquiring a play page corresponding to the user identifier comprises the following steps: A play type corresponding to the user identifier is acquired; When the play type is not played, an initial play page is used as the play page corresponding to the user identifier, and the initial play page is used to play a plurality of sets of explanation videos in a preset order; When the play type is interrupted, a last time interrupted play page is used as the play page corresponding to the user identifier, and the interrupted play page is used to play a set of explanation videos corresponding to the interrupted play page and a set of explanation videos after the set of explanation videos in the preset order; When the play type is ended, an end play page is used as the play page corresponding to the user identifier.

4. The virtual object interaction method of claim 3, wherein, The process of obtaining a plurality of the explanation video groups comprises: obtaining device configuration information and network configuration information corresponding to the terminal device; obtaining a definition type corresponding to the device configuration information and the network configuration information; obtaining a plurality of the explanation video groups corresponding to the definition type.

5. The virtual object interaction method of claim 3, wherein, Each of the playing pages is provided with a side navigation area and / or a bottom navigation area; The side navigation area is provided with one or more of the following function buttons: catalog, subscription, sharing, form, trial, cooperation and external link jump; The bottom navigation area is provided with one or more of the following function buttons: pause / resume, catalog, mode switching, hold to speak and question.

6. The virtual object interaction method of claim 5, wherein, The bottom navigation area is provided with a question button; The terminal device displays N preset questions in the form of a floating layer.

7. The virtual object interaction method of claim 1, wherein, The method further comprises: in response to a click operation on the question button, displaying a text input box and a hold to speak button in the bottom navigation area; in response to a text input operation on the text input box, storing a personalized question corresponding to the most recent video node and the text input operation into a preset message queue; in response to a long press operation on the hold to speak button, converting speech corresponding to the long press operation into text content, and storing the personalized question corresponding to the most recent video node and the converted speech into the preset message queue; when the number of messages in the preset message queue is greater than a preset number, consuming all the messages in the preset message queue to update model parameters of the question recommendation model using the messages in the preset message queue, each of the messages comprising a pair of associated video node and personalized question.

8. The virtual object interaction method of claim 7, wherein, The process of updating the model parameters of the question recommendation model using the messages in the preset message queue comprises: for each of the messages in the preset message queue, performing the following processing: obtaining a first sentiment type corresponding to the personalized question in the message; inputting the video node in the message and the first sentiment type into the question recommendation model to obtain N predicted questions; updating the model parameters of the question recommendation model based on the personalized question in the message and the N predicted questions.

9. The virtual object interaction method of claim 8, wherein, The updating of the model parameters of the question recommendation model based on the personalized question in the message and the N predicted questions comprises: respectively calculating the similarity between the personalized question in the message and the N predicted questions; calculating a prediction error between the personalized question in the message and the predicted question with the lowest similarity; when the prediction error is greater than a preset error, updating the model parameters of the question recommendation model; The process of updating the model parameters of the question recommendation model using the messages in the preset message queue further comprises: when the prediction error is not greater than the preset error, not updating the model parameters of the question recommendation model.

10. An electronic device, comprising: The electronic device comprises a memory and at least one processor, the memory stores a computer program, and the at least one processor implements the following steps when executing the computer program: In response to an access request of a terminal device, a waiting page is displayed by using the terminal device in the process of establishing a connection between the terminal device and a target server, the access request is used to indicate a user identifier; After the connection is established, a guide page is displayed by using the terminal device, and the guide page is provided with a guide button; In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by using the terminal device; Each of the play pages is provided with a question button; In response to a click operation on the question button, a set of explanation videos being played by the terminal device is acquired; Based on the set of explanation videos being played by the terminal device, N preset questions corresponding to the set of explanation videos are acquired, N being a positive integer; In response to a click operation on one of the preset questions, a set of explanation videos corresponding to the clicked preset question is acquired, wherein each set of explanation videos is provided with one or more video nodes; The set of explanation videos corresponding to the clicked preset question is played by using the terminal device; The acquiring of the N preset questions corresponding to the set of explanation videos being played by the terminal device based on the set of explanation videos comprises acquiring a latest video node based on a playing progress of the set of explanation videos being played by the terminal device; N preset questions corresponding to the latest video node are acquired; The acquiring of the N preset questions corresponding to the latest video node comprises acquiring the N preset questions corresponding to the latest video node by using a question recommendation model; The acquiring of the N preset questions corresponding to the latest video node by using the question recommendation model comprises acquiring a second emotional type corresponding to the user identifier; The latest video node and the second emotional type are input into the question recommendation model to obtain the N preset questions; The acquiring of the second emotional type corresponding to the user identifier comprises: A user image is acquired by using an image acquisition component of the terminal device, and an emotional type corresponding to the user image is recognized as the second emotional type corresponding to the user identifier.

11. A computer readable storage medium characterized by, The computer readable storage medium stores a computer program, and the computer program is executed by at least one processor to implement the following steps: In response to an access request of a terminal device, a waiting page is displayed by using the terminal device in the process of establishing a connection between the terminal device and a target server, the access request is used to indicate a user identifier; After the connection is established, a guide page is displayed by using the terminal device, and the guide page is provided with a guide button; In response to a click operation on the guide button, a play page corresponding to the user identifier is displayed by using the terminal device; Each of the play pages is provided with a question button; In response to a click operation on the question button, a set of explanation videos being played by the terminal device is acquired; Based on the terminal device is playing the explanation video group, obtain the corresponding N preset questions, N is positive integer; In response to a click operation for one of the preset questions, obtain the explanation video group corresponding to the clicked preset question; wherein each explanation video group is provided with one or more video nodes; Play the explanation video group corresponding to the clicked preset question by using the terminal device; Based on the terminal device is playing the explanation video group, obtain the corresponding N preset questions, including: based on the terminal device is playing the explanation video group playing progress, obtain the latest one video node; Obtain N preset questions corresponding to the latest one video node; The latest one video node corresponding to the N preset questions, including: using a question recommendation model to obtain the latest one video node corresponding to the N preset questions; The latest one video node corresponding to the N preset questions, including: using a question recommendation model to obtain the latest one video node corresponding to the N preset questions; The latest one video node and the second emotion type are input into the question recommendation model to obtain N preset questions; The latest one video node and the second emotion type are input into the question recommendation model to obtain N preset questions; The latest one video node and the second emotion type are input into the question recommendation model to obtain N preset questions; The terminal device is used to collect user images by using an image acquisition component; the emotion type corresponding to the user image is identified as the second emotion type corresponding to the user identifier. The terminal device is used to collect user images by using an image acquisition component; the emotion type corresponding to the user image is identified as the second emotion type corresponding to the user identifier.

Citation Information

Patent Citations

  • Video information display method and device, storage medium and electronic equipment

    CN111510760A

  • Live broadcast question and answer and interface display method and computer storage medium

    CN114430490A