Virtual object interaction method, electronic equipment, storage medium and program product

By building multiple video queues and preset push sequences, and generating Q&A videos in response to user inquiries, the problem of poor user interaction experience in existing live broadcast technologies is solved, and higher interactivity and user experience is achieved.

CN119996768APending Publication Date: 2025-05-13QUNAR COM BEIJING INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510126739.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-27
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the existing live broadcast technology, the user's interaction experience is poor, and there are problems such as video lag, unnatural expression, poor spoken speech, distortion of sound, and improper connection during live broadcast scene switching.

Method used

By building a first video queue containing multiple script videos and a second video queue containing question-and-answer videos, and pushing the videos in a preset order, generating question-and-answer videos in response to user inquiries, improving the interactivity and user experience of live broadcasts.

Benefits of technology

It improves the interactivity and user experience of live broadcasts, enhances the attractiveness and user participation of live broadcasts, and solves the problem of poor user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996768A_ABST
    Figure CN119996768A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual object interaction method, electronic equipment, a storage medium and a program product. The method comprises the following steps: pushing at least one script video in a first video queue according to a first preset pushing sequence; in response to a received inquiry instruction of the client for a target script video in the at least one script video, generating a question and answer video based on inquiry information corresponding to the inquiry instruction, the question and answer video being used for representing a video in which the virtual object answers the inquiry information; and after adding the question and answer videos into a second video queue, pushing the question and answer videos according to a second preset pushing sequence. The technical problem that the user interaction experience feeling is poor in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a virtual object interaction method, electronic equipment, storage medium and program product. Background Art

[0002] With the development of live broadcast technology and the increase in user demand, the online travel agency (OTA) industry's traditional live broadcast method using digital humans can no longer meet users' high requirements for interactivity and experience quality.

[0003] At present, most live broadcast solutions are limited to the script broadcast and product explanation of digital human hosts, lacking effective user interaction, and there are problems such as video freezes, unnatural expressions of digital humans, poor spoken language conversion, sound distortion, and improper connection when switching live broadcast scenes. These problems not only limit the attractiveness of live broadcasts, but also affect the user interaction experience, resulting in a poor user interaction experience in related technologies.

[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention

[0005] The embodiments of the present invention provide a virtual object interaction method, an electronic device, a storage medium and a program product, so as to at least solve the technical problem of poor user interaction experience in the related art.

[0006] According to one aspect of an embodiment of the present invention, a virtual object interaction method is provided, comprising: pushing at least one script video in a first video queue according to a first preset push order; in response to receiving an inquiry instruction from a client for a target script video in at least one script video, generating a question-and-answer video based on inquiry information corresponding to the inquiry instruction, wherein the question-and-answer video is used to represent a video of the virtual object answering the inquiry information; after adding the question-and-answer video to a second video queue, pushing the question-and-answer video according to a second preset push order.

[0007] Furthermore, based on the inquiry information corresponding to the inquiry instruction, a question-and-answer video is generated, including: based on the inquiry information corresponding to the inquiry instruction, determining target interaction information corresponding to the inquiry information from multiple interaction information, wherein the target interaction information includes preset inquiry information and reply information corresponding to the preset inquiry information, and the similarity between the preset inquiry information and the inquiry information in the target interaction information is greater than a preset threshold; and rendering a virtual object based on the target interaction information to generate a question-and-answer video.

[0008] Further, determining target interaction information corresponding to the query information from multiple interaction information includes: determining multiple similarities between the query information and the multiple interaction information; in response to the presence of a first target similarity greater than or equal to a first threshold among the multiple similarities, determining the interaction information corresponding to the first target similarity among the multiple interaction information as the target interaction information; in response to the absence of the first target similarity greater than the first threshold among the multiple similarities, determining the target interaction information based on a second threshold and the multiple similarities, wherein the second threshold is less than the first threshold.

[0009] Furthermore, target interaction information is determined based on a second threshold and multiple similarities, including: in response to the presence of a second target similarity greater than or equal to the second threshold among the multiple similarities, determining that interaction information corresponding to the second target similarity in the multiple interaction information is preset interaction information; inputting the query information and the preset interaction information into a language processing model, and using the language processing model to adjust the preset interaction information to obtain the target interaction information.

[0010] Furthermore, the method also includes: in response to receiving the script information, rendering the virtual object based on the script information to generate a script video; and adding the script video to the first video queue.

[0011] Further, the method also includes: adjusting the first video stream in the first video queue based on the first preset parameter to obtain a first adjusted video stream, wherein the first preset parameter includes at least one of the following: a first display parameter, a first key frame parameter, and a first transmission parameter, the first display parameter is used to indicate the parameter for adjusting the display effect of the first video stream, the first key frame parameter is used to indicate the parameter for adjusting the interval between multiple key frames in the first video stream, and the first transmission parameter is used to indicate the parameter for adjusting the transmission path of the first video stream; adjusting the second video stream in the second video queue based on the second preset parameter to obtain a second adjusted video stream, wherein the second preset parameter includes at least one of the following: a second display parameter, a second key frame parameter, and a second transmission parameter, the second display parameter is used to indicate the parameter for adjusting the display effect of the second video stream, the second key frame parameter is used to indicate the parameter for adjusting the interval between multiple key frames in the second video stream, and the second transmission parameter is used to adjust the transmission path of the second video stream; pushing based on the first adjusted video stream in the first video queue; pushing based on the second adjusted video stream in the second video queue.

[0012] Furthermore, the method also includes: determining the client to be pushed, and obtaining the first video format of the first video stream, the second video format of the second video stream, the network status of the streaming media server, and the client parameters of the client to be pushed, wherein the streaming media server is used to represent the server that pushes the first video stream and the second video stream; based on the first video format, the network status, and the client parameters, determining the first preset parameters; based on the second video format, the network status, and the client parameters, determining the second preset parameters.

[0013] Furthermore, determining the client to be pushed includes: obtaining the access volume of multiple cache servers, wherein different cache servers are located in different areas, and the multiple cache servers are communicatively connected with the streaming media server; determining a target cache server among the multiple cache servers based on the access volume, wherein the access volume of the target cache server is greater than the access volume of other cache servers among the multiple cache servers except the target cache server; and determining that the client accessing the target cache server is the client to be pushed.

[0014] Furthermore, the method also includes: identifying the script video in the first adjustment video stream to obtain first identification information, and displaying the first identification information during the playback of the script video; identifying the question and answer video in the second adjustment video stream to obtain second identification information, and displaying the second identification information during the playback of the question and answer video.

[0015] According to another aspect of an embodiment of the present invention, a virtual object interaction device is also provided, including: a first push module, used to push at least one script video in a first video queue according to a first preset push order; a generation module, used to generate a question and answer video in response to receiving an inquiry instruction from a client for a target script video in at least one script video, based on inquiry information corresponding to the inquiry instruction, wherein the question and answer video is used to represent a video of the virtual object answering the inquiry information; a second push module, used to push the question and answer video according to a second preset push order after adding the question and answer video to a second video queue.

[0016] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the above-mentioned virtual object interaction method is executed when the program is running.

[0017] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is further provided, the computer-readable storage medium including a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned virtual object interaction method.

[0018] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the computer program implements the above-mentioned virtual object interaction method.

[0019] In an embodiment of the present invention, first, multiple script videos in the first video queue are pushed according to a first preset push order; when a user inquires about the target script video being played through a client, the system responds to the received inquiry instruction, obtains the inquiry information in the inquiry instruction, and generates a question-and-answer video based on the inquiry information; finally, the question-and-answer video is placed in the second video queue, and the question-and-answer video is pushed according to the second preset push order. It is easy to notice that the present application constructs a first video queue containing multiple script videos and a second video queue containing question-and-answer videos, and pushes the script videos in the first video queue and the question-and-answer videos in the second video queue respectively according to the first preset push order and the second preset push order, and at the same time generates an answer video based on the inquiry information in the inquiry instruction, and answers the inquiry information through the virtual object in the question-and-answer video, thereby achieving the technical purpose of improving the live broadcast effect and improving the user interaction experience. In this process, the coherence of the live broadcast content is guaranteed, the interactivity and user experience of the live broadcast are enhanced, thereby improving the user's sense of participation and satisfaction, and thus solving the technical problem of poor user interaction experience in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0021] Figure 1 is a flow chart of a virtual object interaction method according to an embodiment of the present invention;

[0022] Figure 2 is a schematic diagram of pushing a dual video queue according to an embodiment of the present invention;

[0023] Figure 3 is a schematic diagram of the architecture of a live broadcasting platform according to an embodiment of the present invention;

[0024] Figure 4 is a schematic diagram of a key frame interval adjustment process of a video stream according to an embodiment of the present invention;

[0025] Figure 5 is a flow chart of an optional virtual object interaction method according to an embodiment of the present invention;

[0026] Figure 6 is a structural schematic diagram of a virtual object interaction method according to an embodiment of the present invention;

[0027] Figure 7 is a schematic diagram of a live broadcast interactive interface according to an embodiment of the present invention;

[0028] Figure 8 is a schematic diagram of a virtual object interaction device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0031] According to an embodiment of the present invention, an embodiment of a method for interacting with a virtual object is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0032] Figure 1 is a flow chart of a virtual object interaction method according to an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0033] Step S102: Push at least one script video in the first video queue according to a first preset push order.

[0034] The above-mentioned first preset push order may refer to a preset plan for the playback and push order of the script video before the live broadcast begins or when no user asks questions in real time. The first preset push order may be used to guide the system on how to manage and play video content. When there is no specific user interaction, the system will follow this order to automatically advance the live broadcast process to ensure the continuity and integrity of the live broadcast. The first preset push order may be determined based on the order in which the script videos are placed in the first video queue, or may be determined based on the logic of the script video, the continuity of the content and the push strategy, which is not limited here.

[0035] The above-mentioned first video queue may refer to a queue containing all script videos to be played. The first video queue provides the system with a structure for organizing and managing script videos, so that the system can play videos in a preset order. At the same time, when a user asks a question, it can quickly interrupt the currently playing video and switch to the second video queue to process the user's question.

[0036] The above-mentioned script video may refer to video content generated according to the script information. The types of script videos may include but are not limited to script videos based on product introductions, script videos based on product recommendations, script videos based on event notifications, etc. The specific script video should be determined based on the script information and is not limited here. The script video can be used to display preset content related to hotels or other businesses, enhance the attractiveness of live broadcasts, and at the same time ensure the continuity of live broadcasts and the audience's viewing experience when users ask fewer questions.

[0037] The above-mentioned push may refer to the process of transmitting video content from the live broadcast platform to the streaming media server, where the live broadcast platform refers to a management platform used to manage and coordinate multiple technical components and services in the live broadcast process. The push mechanism can be used to ensure that the video content can be transmitted to users with high quality and low latency. Whether it is a script video or a question-and-answer video, it can smoothly reach the user's client and provide users with a good viewing experience.

[0038] In an optional embodiment, the combination of the first preset push order and the first video queue provides structured video content management for live broadcast. In this process, by placing multiple script videos in the first video queue and setting the first preset push order, the live broadcast center pushes multiple script videos according to the first preset push order, ensuring that these script videos can be transmitted to the client in a timely and efficient manner.

[0039] Step S104, in response to receiving a query instruction from the client for a target script video in at least one script video, generating a question-and-answer video based on query information corresponding to the query instruction, wherein the question-and-answer video is used to represent a video of a virtual object answering the query information.

[0040] The above-mentioned client may refer to a software application or web page used by users to watch and interact with live broadcasts. The types of clients may include but are not limited to mobile clients, desktop clients, and web clients. The specific client needs to be determined based on the actual device used by the user. The client can be used to receive and display script videos and question-and-answer videos, and at the same time support users to input commands or ask questions, and interact with the live broadcast system in real time, thereby enhancing user participation and experience.

[0041] The above-mentioned target script video may refer to the script video that is being played or has been played recently when the user asks a question during the live broadcast. The target script video is the context or background of the user's inquiry, which can help the system understand at which time point in which video the user's question was asked.

[0042] The above-mentioned inquiry instruction may refer to an instruction sent by a user to the live broadcast system through a client. The inquiry instruction may be used to ask questions or request information. The inquiry instruction usually includes the user's inquiry information, question timestamp and other information so that the system can understand and respond to the user's inquiry. The inquiry instruction may serve as a link for the user to interact with the live broadcast system. It triggers the generation process of the question-and-answer video, which helps to improve the interactivity and personalization of the live broadcast.

[0043] The above-mentioned inquiry information may refer to the specific questions or request content contained in the inquiry instruction, which is the part that the user hopes to obtain answers or information. The inquiry information is the core input in the question-and-answer video generation process. The system clarifies the user's intention based on the inquiry information, determines the interaction information, and thus generates the most relevant question-and-answer video.

[0044] The above-mentioned question-and-answer video may refer to a question-and-answer video, which is a video clip generated based on user inquiry information through artificial intelligence and other technologies, and contains virtual objects answering questions. The question-and-answer video can be used to directly respond to user inquiries and provide users with the information they need, while increasing the interactivity and fun of the live broadcast and improving the user experience.

[0045] The above-mentioned virtual objects may refer to digital people representing hotels, brands or specific roles in live broadcasts, and are the main form of presentation of live broadcasts. The types of virtual objects may include but are not limited to virtual human-type virtual objects, cartoon image virtual objects, static image-type virtual objects, etc. The specific virtual images can be determined based on system design and user needs, and are not limited here. Virtual objects can be used to interact with users in a visual form, provide information, answer questions, display products, etc., to enhance the attractiveness and interactivity of live broadcasts.

[0046] In an optional embodiment, the user sends an inquiry instruction through the client based on a target script video in multiple script videos, and the live broadcast system generates a question-and-answer video matching the user's inquiry information based on the user's inquiry information in the inquiry instruction, and the virtual object in the question-and-answer video can answer the user's inquiry information in a natural and colloquial way. This process not only reflects the timeliness of the virtual object interaction method proposed in this application in processing user inquiry information, and demonstrates the system's efficient processing ability for inquiry information, but also enriches the content of the live broadcast, improves the interactivity and user experience of the live broadcast, and reflects the technical innovation and practicality of the live broadcast system in this application.

[0047] Step S106, after adding the question and answer video to the second video queue, the question and answer video is pushed according to the second preset push order.

[0048] The above-mentioned second video queue may refer to a queue for storing and managing question-and-answer videos. When a user asks a question and the system generates a question-and-answer video, these videos will be temporarily stored in this queue, waiting to be pushed to the streaming media server. The types of the second video queue may include but are not limited to first-in-first-out queues, priority queues, and local memory queues. The specific type of the second video queue needs to be determined according to actual needs and is not limited here. The second video queue can be used as a temporary storage for question-and-answer videos, so that the system can dynamically manage the generation and push of question-and-answer videos, ensuring that when users ask questions, they can respond quickly and play the question-and-answer video instantly, thereby improving the interactivity and user experience of live broadcasts.

[0049] The above-mentioned second preset push order may refer to a pre-set order for pushing question and answer videos from the second video queue to the streaming media server. The types of the second preset order may include but are not limited to timestamp priority order, user priority order, and question importance order, etc. The specific second preset order needs to be determined according to actual conditions and is not limited here. The second preset push order can be used to ensure that the question and answer video can be quickly pushed to the streaming media server according to a better strategy, so that it can be immediately displayed to users during the live broadcast, thereby maintaining the continuity and immediacy of the live broadcast.

[0050] In an optional embodiment, Figure 2 is a schematic diagram of pushing a dual video queue according to an embodiment of the present invention, such as Figure 2 As shown, Figure 2 The first video queue and the second video queue are input into the intermediate link for processing, and then the streaming media server in the intermediate link pushes the first video queue and the second video queue to the client. The combination of the second video queue and the second preset push sequence can effectively improve the efficiency of processing user questions in live broadcast and the quality of live broadcast interaction.

[0051] Specifically, when the live broadcast station receives a user's question and generates a question-and-answer video, it will add it to the second video queue and, according to the second preset push order, push the question-and-answer video to the streaming media server first, interrupting the currently playing script video and playing the question-and-answer video immediately to quickly respond to the user's question. This mechanism ensures the interactivity and dynamism of the live broadcast, improves the ability to handle a large number of user questions during the live broadcast, and enhances the attractiveness of the live broadcast and the user experience.

[0052] In an optional embodiment, Figure 3 is a schematic diagram of the architecture of a live broadcasting platform according to an embodiment of the present invention. Figure 3 As shown, the live broadcast platform mainly includes the access layer, core application layer, infrastructure layer, business scenarios and management modules. Among them, the access layer includes interface protocols and live broadcast streams, the core application layer includes live broadcast room management, live broadcast configuration, content management, streaming server management and streaming core scheduling, the infrastructure layer includes streaming core clusters and streaming media server clusters, and management includes user management and other management.

[0053] Specifically, the access layer of the live broadcast platform supports the access of multiple protocols and the processing of multi-protocol live streams; the core application layer includes live broadcast room management, live broadcast configuration, content management, stream server management and risk control strategies, and supports the creation, configuration, content maintenance and stream management of live broadcasts; the infrastructure layer includes the streaming core cluster and the streaming media server cluster, which are responsible for the processing and transmission of audio and video streams. The live broadcast platform completes the push of video streams and live broadcast management through the call of the above modules.

[0054] In an embodiment of the present invention, first, multiple script videos in the first video queue are pushed according to a first preset push order; when a user inquires about the target script video being played through a client, the system responds to the received inquiry instruction, obtains the inquiry information in the inquiry instruction, and generates a question-and-answer video based on the inquiry information; finally, the question-and-answer video is placed in the second video queue, and the question-and-answer video is pushed according to the second preset push order. It is easy to notice that the present application constructs a first video queue containing multiple script videos and a second video queue containing question-and-answer videos, and pushes the script videos in the first video queue and the question-and-answer videos in the second video queue respectively according to the first preset push order and the second preset push order, and at the same time generates an answer video based on the inquiry information in the inquiry instruction, and answers the inquiry information through the virtual object in the question-and-answer video, thereby achieving the technical purpose of improving the live broadcast effect and improving the user interaction experience. In this process, the coherence of the live broadcast content is guaranteed, the interactivity and user experience of the live broadcast are enhanced, thereby improving the user's sense of participation and satisfaction, and thus solving the technical problem of poor user interaction experience in the related technology.

[0055] Optionally, based on the inquiry information corresponding to the inquiry instruction, a question and answer video is generated, including: based on the inquiry information corresponding to the inquiry instruction, determining target interaction information corresponding to the inquiry information from multiple interaction information, wherein the target interaction information includes preset inquiry information and reply information corresponding to the preset inquiry information, and the similarity between the preset inquiry information and the inquiry information in the target interaction information is greater than a preset threshold; and rendering a virtual object based on the target interaction information to generate a question and answer video.

[0056] The above-mentioned multiple interactive information may refer to multiple question-and-answer information stored in the knowledge base for answering user inquiry information. The interactive information includes user inquiry information and reply information corresponding to the inquiry information. The multiple interactive information can be used to provide standard reply information for user inquiry information, simplifying the process of directly generating reply information based on inquiry information and improving the efficiency of replying to inquiry information.

[0057] The above-mentioned target interaction information may refer to one or a group of interaction information that is most relevant to the user's current inquiry and is screened out from multiple interaction information. The target interaction information can be used to ensure that the question-and-answer video content generated by the system is highly relevant to the user's inquiry, thereby improving the accuracy and timeliness of the question-and-answer, and enhancing the interactivity and user satisfaction of the live broadcast.

[0058] The above-mentioned preset inquiry information may refer to the user's inquiry information contained in the target interactive information. The types of preset inquiry information may include but are not limited to hotel reservation questions, facility service questions, and price policy questions. The specific type of preset inquiry information needs to be determined based on the question classification and inquiry information content. No limitation is made here. The preset inquiry information can be used to enable the system to quickly respond to user questions, and by matching relatively similar preset questions, corresponding question-and-answer videos can be quickly generated to improve the interactivity and smoothness of the live broadcast.

[0059] The above-mentioned similarity may refer to the degree of similarity between the inquiry information and the preset inquiry information. The similarity calculation method may include but is not limited to cosine similarity, Euclidean distance, edit distance, etc. The specific similarity calculation method needs to be determined according to the actual situation and is not limited here. The similarity calculation can be used to determine the matching degree between the user inquiry and the preset inquiry information, which is the key basis for the selection of target interaction information, ensuring the accuracy and pertinence of the question and answer video content.

[0060] The above-mentioned preset threshold may refer to a set similarity standard, which is used to determine whether the inquiry information and the preset inquiry information in the target interaction information are similar enough to trigger the generation of a question-and-answer video. The preset threshold may be determined based on the requirements for interaction accuracy and system performance, and is not limited here. The preset threshold may be used to ensure that the system does not generate a question-and-answer video when the similarity between the inquiry information and the preset inquiry information is low, thereby avoiding the playback of irrelevant or erroneous content and improving user experience and interaction quality.

[0061] The above-mentioned rendering may refer to the process of converting the answer information in the target interactive information into a visual video. The types of rendering may include but are not limited to real-time rendering, cloud rendering, etc. The specific rendering method needs to be determined according to actual needs and system design, and is not limited here. The rendering technology converts the answer information into an intuitive question-and-answer video, so that the virtual object can vividly display the answer process, thereby enhancing the visual appeal and interactivity of the live broadcast.

[0062] In an optional embodiment, when the client sends an inquiry instruction, the system accurately locks the target interaction information that matches the user's question from a large amount of pre-stored interaction information. This process calculates the similarity between the preset inquiry information and the user's inquiry information. When the similarity between the two exceeds the preset threshold, that is, it reaches the sufficient relevance standard set by the system, the target interaction information will be determined; then the system uses cloud rendering technology to integrate and render the target interaction information with the virtual object to generate a question-and-answer video, in which the virtual object presents the answer in a natural and smooth manner, greatly improving the experience and efficiency of live interaction. In this way, the present application significantly enhances the user participation in the live broadcast of digital humans, ensures the accuracy and timeliness of the live content, and at the same time achieves high experience, high fluency and high interactivity of the live content, greatly improving the competitiveness and user satisfaction of the live platform.

[0063] Optionally, determining target interaction information corresponding to the query information from multiple interaction information includes: determining multiple similarities between the query information and the multiple interaction information; in response to the presence of a first target similarity greater than or equal to a first threshold among the multiple similarities, determining the interaction information corresponding to the first target similarity among the multiple interaction information as the target interaction information; in response to the absence of the first target similarity greater than the first threshold among the multiple similarities, determining the target interaction information based on a second threshold and the multiple similarities, wherein the second threshold is less than the first threshold.

[0064] The above-mentioned first threshold may refer to a higher similarity threshold that is set. When there is a similarity between the query information and multiple interaction information that exceeds the first threshold, the similarity is determined to be the first target similarity, and the corresponding interaction information is the target interaction information. The first threshold may include but is not limited to 95%, 93% and 90%, etc. The specific first threshold can be determined based on the system processing capability and the size of the knowledge base, and is not limited here. The setting of the first threshold can be used to directly determine the interaction information that is most relevant to the query information, which simplifies the user intent recognition step and improves the efficiency of question and answer video generation.

[0065] The first target similarity may refer to the maximum similarity value among multiple similarities that reaches or exceeds the first threshold. The first target similarity is the direct basis for the system to determine the target interaction information, and represents the highest similarity between the user query information and a preset interaction information.

[0066] The above-mentioned second threshold may refer to a lower standard threshold for further screening interactive information when the first threshold screening fails. When there is no preset information with a similarity to the user's query that reaches the first threshold, the system will consider the second threshold to find relatively matching interactive information. The second threshold may include but is not limited to 88%, 86%, and 85%, etc. The specific second threshold can be determined based on the system processing capacity and the size of the knowledge base, which is not limited here. The existence of the second threshold ensures that even in the absence of highly matching information, the system can try its best to respond to the user, avoiding the possible omission of useful information due to the strict first threshold standard, and improving the coverage of interaction and the system's ability to handle complex problems.

[0067] In an optional embodiment, when there is a similarity between the query information and multiple interactive information that exceeds a first threshold, the similarity is determined to be the first target similarity, and the interactive information corresponding thereto is the target interactive information; when there is no similarity between the query information and multiple interactive information that does not exceed the first threshold, the target interactive information is further determined based on the second threshold. In the above process, by setting the first threshold and the second threshold, the system can flexibly and accurately handle user questions, which not only ensures the accuracy of the answers, but also takes into account the wide coverage of the questions, significantly improving the experience and efficiency of live interactive.

[0068] Optionally, target interaction information is determined based on a second threshold and multiple similarities, including: in response to the presence of a second target similarity greater than or equal to the second threshold among the multiple similarities, determining that interaction information corresponding to the second target similarity in the multiple interaction information is preset interaction information; inputting the query information and the preset interaction information into a language processing model, and using the language processing model to adjust the preset interaction information to obtain the target interaction information.

[0069] The above-mentioned second target similarity may refer to the maximum similarity value among multiple similarities that reaches or exceeds the second threshold, and is used to find relatively matching interaction information when there is no highly matching preset query information. The setting of the second target similarity ensures that when the system faces user questions, it will not miss possible relevant information due to the strict first threshold. Through more relaxed standards, the coverage of questions and answers and the response rate of user questions are improved.

[0070] The above-mentioned preset interaction information may refer to question and answer information that is determined to be relatively matched with the user's inquiry information after similarity calculation and second threshold screening. The preset interaction information is the basis for the system to generate question and answer videos. It contains standard questions and corresponding answers prepared by the system, and is the original material for adjusting the language processing model.

[0071] The above-mentioned language processing model may refer to an artificial intelligence model. The types of language processing models may include but are not limited to rule-based language processing models, statistics-based language processing models, or deep learning-based large language processing models (Large Language Model, LLM for short), etc. The specific language processing model can be determined according to the system design requirements and is not limited here. The language processing model can make the answer more natural and personalized by adjusting the preset interaction information, thereby improving the user's satisfaction with questions and the live interactive experience.

[0072] The above-mentioned adjustment may refer to polishing and modifying the standard answers in the preset interactive information based on the specific content and context of the user's question, so as to generate a question-and-answer video text that is more in line with user needs and spoken. The adjustment methods may include but are not limited to colloquial polishing, user personalized adjustment, contextual adaptability adjustment, etc. The specific adjustment strategy needs to be determined based on the capabilities of the language processing model and the design of the live broadcast system. The adjustment process ensures that the content of the question-and-answer video is not only accurate, but also can be presented in a more natural and daily communication manner, thereby improving user acceptance of the answers and the interactivity of the live broadcast.

[0073] In an optional embodiment, the similarity is first calculated based on the user query information and the preset interaction information, and the relatively matching preset interaction information is determined through the second target similarity; then the user query information and the preset interaction information are input into the language processing model for adjustment to generate a more colloquial Q&A video text that is closer to user needs. This series of operations not only ensures that user questions can be responded to in a timely manner, but also improves the quality of the Q&A video and enhances the interactivity and user experience of the live broadcast.

[0074] Optionally, the method further includes: in response to receiving the script information, rendering the virtual object based on the script information to generate a script video; and adding the script video to the first video queue.

[0075] The above-mentioned script information may refer to a pre-prepared script or text content, which is used to guide the behavior, language and expression of the virtual object, and ensure that the information output during the live broadcast conforms to the expected process and theme. The type of script information may include but is not limited to hotel facility introduction, service features, booking process instructions, promotional activity announcements, etc. The specific script information can be determined according to actual conditions and is not limited here.

[0076] In an optional embodiment, in response to receiving the script information, the system uses rendering technology to combine the script content with virtual objects to generate script videos. These script videos cover information on hotel facilities, service features, booking processes, promotional activities, etc., which not only enriches the live broadcast content, but also ensures that the live broadcast process remains coherent, professional and attractive when there are no immediate user questions. After the script video is generated, it will be added to the first video queue, which is the core component of the system and is responsible for managing the video stream to be played. When the live broadcast center detects that there is no user question, it will take out the script video from the first video queue for playback. Once a user asks a question, the system will quickly switch to the playback of the question-and-answer video, achieving seamless connection and flexible scheduling of the live broadcast content.

[0077] Optionally, the method also includes: adjusting the first video stream in the first video queue based on first preset parameters to obtain a first adjusted video stream, wherein the first preset parameters include at least one of the following: a first display parameter, a first key frame parameter, and a first transmission parameter, the first display parameter is used to indicate a parameter for adjusting a display effect of the first video stream, the first key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the first video stream, and the first transmission parameter is used to indicate a parameter for adjusting a transmission path of the first video stream; adjusting the second video stream in the second video queue based on second preset parameters to obtain a second adjusted video stream, wherein the second preset parameters include at least one of the following: a second display parameter, a second key frame parameter, and a second transmission parameter, the second display parameter is used to indicate a parameter for adjusting a display effect of the second video stream, the second key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the second video stream, and the second transmission parameter is used to indicate a parameter for adjusting a transmission path of the second video stream; pushing based on the first adjusted video stream in the first video queue; pushing based on the second adjusted video stream in the second video queue.

[0078] The above-mentioned first preset parameters may refer to pre-set parameters for adjusting the display effect, key frame interval and transmission path of the first video stream. The first preset parameters may include but are not limited to first display parameters, first key frame parameters, first transmission parameters, etc. The specific first preset parameters can be adjusted according to actual needs and are not limited here. The first preset parameters can dynamically adjust the properties of the video stream according to different playback videos, thereby improving video quality and user experience, while reducing transmission delays and bandwidth occupancy, and improving the overall performance and efficiency of the system.

[0079] The above-mentioned first video stream may refer to a script video stream.

[0080] The above-mentioned first adjusted video stream may refer to a video stream after the first preset parameter adjustment is applied to the first video stream. The adjusted video stream can better adapt to the playback scenario and network conditions. The first adjusted video stream ensures that users can obtain a smooth and high-quality video playback experience in different viewing environments and network conditions, while reducing the consumption of system resources and network delays, and improving the real-time and reliability of live broadcasts.

[0081] The above-mentioned first display parameter and second display parameter respectively refer to parameters used to adjust the display effects of the first video stream and the second video stream. The first display parameter and the second display parameter may include but are not limited to display parameters such as complex filter parameters. The specific first display parameter and the second display parameter need to be adjusted according to actual needs and are not limited here.

[0082] The above-mentioned first key frame parameters and second key frame parameters respectively refer to parameters used to adjust the key frame intervals of the first video stream and the second video stream. The first key frame parameters and the second key frame parameters may include but are not limited to outputting additional parameters, etc. The specific first key frame parameters and the second key frame parameters need to be adjusted according to actual needs and are not limited here.

[0083] The above-mentioned first transmission parameter and second transmission parameter respectively refer to parameters used to adjust the transmission effects of the first video stream and the second video stream. The first transmission parameter and the second transmission parameter may include but are not limited to transmission protocol selection parameters, etc. The specific first transmission parameter and the second transmission parameter need to be selected according to actual needs and are not limited here.

[0084] The above-mentioned second preset parameters may refer to pre-set parameters for adjusting the display effect, key frame interval and transmission path of the second video stream. The second preset parameters may include but are not limited to second display parameters, second key frame parameters, second transmission parameters, etc. The specific second preset parameters can be adjusted according to actual needs and are not limited here. The second preset parameters can dynamically adjust the properties of the video stream according to different playback videos, thereby improving video quality and user experience, while reducing transmission delays and bandwidth occupancy, and improving the overall performance and efficiency of the system.

[0085] The above-mentioned second video stream may refer to a question-and-answer video stream.

[0086] The above-mentioned second adjusted video stream may refer to a video stream after the second preset parameter adjustment is applied to the second video stream. The adjusted video stream can better adapt to the playback scenario and network conditions. The second adjusted video stream ensures that users can obtain a smooth and high-quality video playback experience in different viewing environments and network conditions, while reducing the consumption of system resources and network delays, and improving the real-time and reliability of live broadcasts.

[0087] In an optional embodiment, the first video stream in the first video queue is adjusted based on the first preset parameter to obtain a first adjusted video stream; the second video stream in the second video queue is adjusted based on the second preset parameter to obtain a second adjusted video stream. The above adjustment process can be performed by adjusting the preset parameters in the video stream through the multimedia processing tool FFmpeg to obtain the adjusted video stream.

[0088] In an optional embodiment, the first video stream and the second video stream are adjusted based on the first preset parameter and the second preset parameter respectively to obtain the first adjusted video stream and the second adjusted video stream, and then the first adjusted video stream and the second adjusted video stream are pushed through the streaming media server. The above process of adjusting different video streams by different preset parameters not only ensures the high quality and smooth playback of the video content, but also enables the system to flexibly switch content according to the real-time interactive situation, thereby improving the interactivity and participation of the live broadcast. At the same time, the dynamic adjustment capability of the preset parameters enables the system to automatically adapt to different devices, network environments and viewing needs, greatly improving the user experience and overall efficiency of the live broadcast, and providing a solid technical foundation for the widespread application of virtual object live broadcast technology.

[0089] In an optional embodiment, Figure 4 is a schematic diagram of a key frame interval adjustment process of a video stream according to an embodiment of the present invention. Figure 4 As shown, Figure 4 The two video streams are the video stream before optimization and the video stream after optimization. Each video stream contains multiple intra-frame coded frames (Intra-frame, referred to as I frames), multiple bidirectional prediction frames (Bi-directional frame, referred to as B frames), and multiple inter-frame prediction frames (Predictive frame, referred to as P frames). The direction of the video stream is from left to right. During the video stream push process, the player pulls the stream at intervals of 1 second. The video stream before optimization contains multiple key frames such as I frames, B frames, and P frames, and these key frames are arranged continuously; the key frames in the optimized video stream are rearranged, and each key frame interval contains an I frame, B frame, and P frame, and the key frame interval becomes shorter. The above process balances the video quality and transmission efficiency by reasonably adjusting the key frame interval, avoiding the code stream, storage and delay problems caused by too many or too few key frames.

[0090] Optionally, the method also includes: determining the client to be pushed, and obtaining the first video format of the first video stream, the second video format of the second video stream, the network status of the streaming media server, and the client parameters of the client to be pushed, wherein the streaming media server is used to represent the server that pushes the first video stream and the second video stream; based on the first video format, the network status, and the client parameters, determining the first preset parameters; based on the second video format, the network status, and the client parameters, determining the second preset parameters.

[0091] The above-mentioned first video format and second video format refer to the encoding and decoding formats and packaging methods of the script video and the question-and-answer video respectively. The types of video formats may include but are not limited to the standard format of Part 4, Item 14 of the Moving Picture Experts Group (MPEG-4 Part 14 format, referred to as MP4 format), Flash Video format (FLV), and Transport Stream format (TS). The types of encoding and decoding formats may include but are not limited to High Efficiency Video Coding (H.265), Video Coding for Video conferencing, Version 9 (VP9), etc. The specific video format and encoding and decoding format need to be determined according to the type of video stream, and are not limited here. Selecting a suitable video format can ensure that the video stream has good compatibility and transmission efficiency while maintaining high quality. According to different types of video content, the system can flexibly adjust the format to meet the needs of different scenarios.

[0092] The above-mentioned streaming media server (Simple-Realtime-Server, referred to as SRS) may refer to a server responsible for processing and pushing video streams to user clients. The streaming media server can support multiple transmission protocols and can intelligently adjust the video stream push strategy according to network status and client parameters to ensure that the video content can be stably and smoothly transmitted to the user device.

[0093] The above network status may refer to the quality of the network connection between the system and the user client. The network status may include but is not limited to indicators such as bandwidth, delay, and packet loss rate. The specific network status indicators need to be determined based on the system design and actual conditions, and are not limited here. By monitoring the network status, the system can adjust the transmission strategy of the video stream in real time, such as encoding parameters, the selection of streaming media protocols, etc., to cope with network fluctuations and maintain the smoothness of video playback.

[0094] The above-mentioned client parameters may refer to specific attributes of the user device. The client parameters may include but are not limited to device model, operating system version, network type, screen size, device storage space, etc. The specific client parameters need to be determined according to actual needs and are not limited here. By obtaining the client parameters, the system can intelligently infer the appropriate video format and encoding parameters to ensure good playback effect of the video stream on the user device, while taking into account the performance limitations of the device to provide an optimized video experience.

[0095] In an optional embodiment, after determining the client to be pushed, the system will obtain the video formats of the first video stream and the second video stream, and monitor the network status of the streaming media server and the device parameters of the client to be pushed; then, the system determines the first preset parameter and the second preset parameter based on this information, which are used to adjust the display effect, key frame interval and transmission path of the script video and the question-and-answer video. This series of operations ensures that the video stream can be pushed to the client in a good state under different playback scenarios and network conditions, providing users with a smooth and high-quality video playback experience, while effectively coping with network fluctuations, reducing video freezes, and improving user satisfaction and the overall effect of live broadcasting.

[0096] Optionally, determining the client to be pushed includes: obtaining the access volume of multiple cache servers, wherein different cache servers are located in different areas, and the multiple cache servers are communicatively connected to the streaming media server; determining a target cache server among the multiple cache servers based on the access volume, wherein the access volume of the target cache server is greater than the access volume of other cache servers among the multiple cache servers except the target cache server; and determining that the client accessing the target cache server is the client to be pushed.

[0097] The above-mentioned cache server may refer to a server deployed in a network for storing and quickly providing frequently accessed data. The type of cache server may be a part of a content delivery network (CDN), which is deployed around the world to cover different geographical areas; or it may be a local cache server, serving a specific network or region. The above-mentioned cache server types are only examples, and the specific cache server needs to be determined according to actual needs and is not limited here. The cache server can be used to reduce the direct burden of the streaming media server and improve the transmission efficiency of the video stream. At the same time, by reducing network latency and data transmission distance, it can improve the smoothness and response speed of users watching live videos.

[0098] The above-mentioned visits may refer to the number of times or the amount of data that a client requests a live video stream from a cache server within a certain period of time. The types of visits may include but are not limited to real-time visits and historical visits. The determination of the visits here needs to be determined based on actual needs and is not limited here. By monitoring and analyzing the visits, the system can understand which cache servers have a higher load and which areas have more intensive user demand, thereby providing data support for optimizing the distribution strategy of the video stream.

[0099] The target cache server may be a cache server with a high visit volume among multiple cache servers. Determining the target cache server can help the system identify which servers carry more user requests, so as to give priority to distributing video streams using these servers, ensuring that most users can get a good video experience.

[0100] In an optional embodiment, Figure 5 is a flow chart of an optional virtual object interaction method according to an embodiment of the present invention. Figure 5 As shown, first, the user sends a question through the client; then the question-answering interactive service recognizes the intent of the query information and matches the interactive information in the knowledge base to obtain the target interactive information, and finally the large language model adjusts and polishes the target interactive information to obtain a colloquial reply; then the live broadcast center calls the text-to-audio technology in the cloud rendering service to convert the text into audio, and then calls the virtual object generation technology to generate a question-and-answer video; then the video stream is adjusted through the multimedia processing tool to obtain the adjusted video stream, and the streaming core of the live broadcast center schedules the streaming media server to push the video stream to the cache server; finally, the cache server pushes it to the client.

[0101] In an optional embodiment, Figure 6 is a structural diagram of a virtual object interaction method according to an embodiment of the present invention. Figure 6 As shown, the method mainly involves the client, question-answer interaction service, live broadcast platform, cloud rendering service, streaming media server and cache server, among which the question-answer interaction service includes intent recognition and spoken reply; the live broadcast platform includes live broadcast room management, live broadcast configuration, content management, user management, streaming media server management, streaming core scheduling, audio and video synthesis adding filters, audio and video encoding, packaging and other functions; the cloud rendering service includes text-to-audio technology and virtual object generation technology.

[0102] In actual applications, the user first sends a question through the client; the question-answering interactive service then processes the query information through intent recognition and spoken replies to obtain the target interactive information; the live broadcast center then calls the text-to-audio technology in the cloud rendering service to convert the text into audio, and then calls the virtual object generation technology to generate a question-and-answer video; the live broadcast center then adjusts the video stream through multimedia processing tools to obtain the adjusted video stream, and the live broadcast center's streaming core scheduling uses the streaming media server to push the video stream to the cache server; finally, the cache server pushes it to the client. In this process, the live broadcast center plays a core scheduling role.

[0103] In an optional embodiment, in the process of determining the client to be pushed, the system first obtains the access volume of multiple cache servers, which are distributed in different geographical areas and maintain communication connection with the streaming media server. Then, based on the access volume data, the system identifies the target cache server with higher access volume. Finally, the system determines that all clients accessing the target cache server are clients to be pushed, which means that these clients will receive the optimized video stream first. Through this mechanism, the system can effectively push the video stream to user groups with more concentrated needs, improve the efficiency of video distribution and user viewing experience, while reducing the load of non-target cache servers and realizing the rational allocation and utilization of resources.

[0104] Optionally, the method also includes: identifying the script video in the first adjustment video stream to obtain first identification information, and displaying the first identification information during the playback of the script video; identifying the question and answer video in the second adjustment video stream to obtain second identification information, and displaying the second identification information during the playback of the question and answer video.

[0105] The above-mentioned first identification information may refer to information related to the video content and which can enhance the user experience obtained by the system through identifying the script video. The types of the first identification information may include but are not limited to links containing product information, interactive reminder information to guide users to participate in interactive activities such as asking questions, voting, and games, and scene information introducing video scenes, etc. The specific first identification information needs to be determined based on the actual content of the script video and is not limited here. The first identification information can be used to help the audience better understand and interact with the script video content, provide real-time feedback and supplementary information, such as product links, hotel reservation details, virtual object role introductions, etc., thereby increasing the interactivity and attractiveness of the live broadcast.

[0106] The above-mentioned second identification information may refer to the information related to the question and answer content obtained by the system through identification and analysis of the question and answer video. The types of the second identification information may include but are not limited to virtual object reply confirmation information, virtual object reply extension information, interactive information, etc. The specific second identification information needs to be determined based on the content of the question and answer video and is not limited here. The second identification information can be used to help the audience understand the content of the virtual object's answer, provide background, explanation or extended information of the question, and enhance the educational and entertainment value of the live broadcast.

[0107] In an optional embodiment, through supplemental enhancement information (SEI), the live broadcast system can intelligently analyze the content of the script video and the question-and-answer video, quickly generate relevant identification information, and seamlessly display it to the audience during the video playback, thereby realizing the intelligence and personalization of the live broadcast content. This not only reflects the technological innovation of the live broadcast system, but also reflects the huge potential of digital human live broadcast in improving user experience and promoting commercial transformation.

[0108] In an optional embodiment, Figure 7 is a schematic diagram of a live interactive interface according to an embodiment of the present invention. Figure 7 As shown, the display content of the live broadcast interface includes: the name of the live broadcast room, virtual objects, product information, dynamic links, a dialogue input box, and pop-up user inquiry information, wherein the user inquiry information may include multiple items, such as user 1 inquiry information, user 2 inquiry information, etc.

[0109] According to another aspect of an embodiment of the present invention, a virtual object interaction device is also provided, which can execute the virtual object interaction method of the above embodiment. The specific implementation method and preferred application scenario are the same as those of the above embodiment and will not be repeated here.

[0110] Figure 8 is a schematic diagram of a virtual object interaction device according to an embodiment of the present invention. Figure 8 As shown, the device includes the following: a first pushing module 802 , a first generating module 804 , and a second pushing module 806 .

[0111] The first push module is used to push at least one script video in the first video queue according to a first preset push order; the first generation module is used to generate a question-and-answer video in response to receiving a query instruction from the client for a target script video in at least one script video, based on the query information corresponding to the query instruction, wherein the question-and-answer video is used to represent a video in which a virtual object answers the query information; the second push module is used to push the question-and-answer video according to a second preset push order after adding the question-and-answer video to the second video queue.

[0112] Optionally, the first generation module includes: a first determination unit, used to determine target interaction information corresponding to the inquiry information from multiple interaction information based on the inquiry information corresponding to the inquiry instruction, wherein the target interaction information includes preset inquiry information and reply information corresponding to the preset inquiry information, and the similarity between the preset inquiry information and the inquiry information in the target interaction information is greater than a preset threshold; and a first generation unit, used to render a virtual object based on the target interaction information to generate a question-and-answer video.

[0113] Optionally, the first determination unit includes: a first determination subunit, used to determine multiple similarities between the query information and multiple interaction information; a second determination subunit, used to determine that the interaction information corresponding to the first target similarity in the multiple interaction information is the target interaction information in response to the presence of a first target similarity greater than or equal to a first threshold among the multiple similarities; and a third determination subunit, used to determine the target interaction information based on the second threshold and multiple similarities in response to the absence of the first target similarity greater than the first threshold among the multiple similarities, wherein the second threshold is smaller than the first threshold.

[0114] Optionally, the third determination subunit includes: in response to the presence of a second target similarity greater than or equal to a second threshold among multiple similarities, determining that the interaction information corresponding to the second target similarity in the multiple interaction information is preset interaction information; inputting the query information and the preset interaction information into the language processing model, and using the language processing model to adjust the preset interaction information to obtain the target interaction information.

[0115] Optionally, the device also includes: a second generating module, used for rendering virtual objects based on the script information in response to receiving the script information, and generating a script video; and adding the script video to the first video queue.

[0116] Optionally, the device further includes: a first adjustment module, used to adjust the first video stream in the first video queue based on a first preset parameter to obtain a first adjusted video stream, wherein the first preset parameter includes at least one of the following: a first display parameter, a first key frame parameter, and a first transmission parameter, the first display parameter is used to indicate a parameter for adjusting a display effect of the first video stream, the first key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the first video stream, and the first transmission parameter is used to indicate a parameter for adjusting a transmission path of the first video stream; a second adjustment module, used to adjust the first video stream in the second video queue based on a second preset parameter The second video stream is adjusted to obtain a second adjusted video stream, wherein the second preset parameters include at least one of the following: a second display parameter, a second key frame parameter, and a second transmission parameter, the second display parameter is used to indicate a parameter for adjusting a display effect of the second video stream, the second key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the second video stream, and the second transmission parameter is used to indicate a parameter for adjusting a transmission path of the second video stream; a third push module is used to push based on the first adjusted video stream in the first video queue; and a fourth push module is used to push based on the second adjusted video stream in the second video queue.

[0117] Optionally, the device also includes: an acquisition module, used to determine the client to be pushed, and obtain the first video format of the first video stream, the second video format of the second video stream, the network status of the streaming media server, and the client parameters of the client to be pushed, wherein the streaming media server is used to represent the server that pushes the first video stream and the second video stream; a first determination module, used to determine the first preset parameters based on the first video format, the network status, and the client parameters; a second determination module, used to determine the second preset parameters based on the second video format, the network status, and the client parameters.

[0118] Optionally, the acquisition module includes: an acquisition unit, used to obtain the access volume of multiple cache servers, wherein different cache servers are located in different areas, and the multiple cache servers are communicatively connected to the streaming media server; a second determination unit, used to determine a target cache server among the multiple cache servers based on the access volume, wherein the access volume of the target cache server is greater than the access volume of other cache servers among the multiple cache servers except the target cache server; and a third determination unit, used to determine that a client accessing the target cache server is a client to be pushed.

[0119] Optionally, the device also includes: a first identification module, used to identify the script video in the first adjustment video stream, obtain first identification information, and display the first identification information during the playback of the script video; a second identification module, used to identify the question and answer video in the second adjustment video stream, obtain second identification information, and display the second identification information during the playback of the question and answer video.

[0120] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the method in each embodiment of the present invention is executed when the program is running.

[0121] According to another aspect of an embodiment of the present invention, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program. When the program is executed, a processor of a device is controlled to execute the methods of various embodiments of the present invention.

[0122] The computer storage medium in the above steps can be a medium used to store certain discontinuous physical quantities in a computer memory, and the computer storage medium mainly includes semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser disks, etc. The stored program included in the computer-readable storage medium can be a set of instructions that can be recognized and executed by a computer, running on an electronic computer, and is an information tool that meets certain needs of people.

[0123] According to another aspect of an embodiment of the present invention, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the method in each embodiment of the present invention is implemented.

[0124] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0125] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0126] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0127] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0128] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.

[0129] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A virtual object interaction method, applied to a server, characterized in that: include: Pushing at least one script video in the first video queue according to a first preset pushing order; In response to receiving a query instruction from a client for a target scenario video in the at least one scenario video, generating a question-and-answer video based on query information corresponding to the query instruction, wherein the question-and-answer video is used to represent a video of a virtual object answering the query information; After the question-and-answer video is added to the second video queue, the question-and-answer video is pushed according to a second preset push order.

2. The method according to claim 1, characterized in that Generating a question-and-answer video based on the inquiry information corresponding to the inquiry instruction includes: Based on the query information corresponding to the query instruction, determine target interaction information corresponding to the query information from multiple interaction information, wherein the target interaction information includes preset query information and reply information corresponding to the preset query information, and the similarity between the preset query information in the target interaction information and the query information is greater than a preset threshold; The virtual object is rendered based on the target interaction information to generate the question-and-answer video.

3. The method according to claim 2, characterized in that Determining target interaction information corresponding to the query information from the plurality of interaction information includes: determining a plurality of similarities between the query information and the plurality of interaction information; In response to a first target similarity being greater than or equal to a first threshold among the multiple similarities, determining the interaction information corresponding to the first target similarity among the multiple interaction information as the target interaction information; In response to the first target similarity not existing in the plurality of similarities that is greater than the first threshold, the target interaction information is determined based on a second threshold and the plurality of similarities, wherein the second threshold is less than the first threshold.

4. The method according to claim 3, characterized in that: Determining the target interaction information based on the second threshold and the multiple similarities includes: In response to a second target similarity being greater than or equal to the second threshold value among the multiple similarities, determining that the interaction information corresponding to the second target similarity among the multiple interaction information is preset interaction information; The query information and the preset interaction information are input into a language processing model, and the preset interaction information is adjusted using the language processing model to obtain the target interaction information.

5. The method according to claim 2, characterized in that: The method further comprises: In response to receiving the script information, rendering the virtual object based on the script information to generate the script video; Add the script video to the first video queue.

6. The method according to claim 1, characterized in that The method further comprises: Adjusting the first video stream in the first video queue based on a first preset parameter to obtain a first adjusted video stream, wherein the first preset parameter includes at least one of the following: a first display parameter, a first key frame parameter, and a first transmission parameter, wherein the first display parameter is used to indicate a parameter for adjusting a display effect of the first video stream, the first key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the first video stream, and the first transmission parameter is used to indicate a parameter for adjusting a transmission path of the first video stream; Adjusting the second video stream in the second video queue based on second preset parameters to obtain a second adjusted video stream, wherein the second preset parameters include at least one of the following: a second display parameter, a second key frame parameter, and a second transmission parameter, the second display parameter is used to indicate a parameter for adjusting a display effect of the second video stream, the second key frame parameter is used to indicate a parameter for adjusting an interval between multiple key frames in the second video stream, and the second transmission parameter is used to indicate a parameter for adjusting a transmission path of the second video stream; Pushing based on the first adjusted video stream in the first video queue; Pushing is performed based on the second adjusted video stream in the second video queue.

7. The method according to claim 6, characterized in that The method further comprises: Determine the client to be pushed, and obtain the first video format of the first video stream, the second video format of the second video stream, the network status of the streaming media server, and the client parameters of the client to be pushed, wherein the streaming media server is used to represent the server that pushes the first video stream and the second video stream; Determining the first preset parameter based on the first video format, the network status, and the client parameter; The second preset parameter is determined based on the second video format, the network status, and the client parameter.

8. The method according to claim 7, characterized in that Determine the client to be pushed, including: Obtaining access volume of multiple cache servers, wherein different cache servers are located in different areas, and the multiple cache servers are communicatively connected with the streaming media server; Determine a target cache server among the multiple cache servers based on the access volume, wherein the access volume of the target cache server is greater than the access volume of other cache servers among the multiple cache servers except the target cache server; The client accessing the target cache server is determined as the client to be pushed.

9. The method according to claim 6, characterized in that The method further includes: Identify the script video in the first adjusted video stream to obtain first identification information, and display the first identification information during the playback of the script video; The question-and-answer video in the second adjusted video stream is identified to obtain second identification information, and the second identification information is displayed during the playback of the question-and-answer video.

10. An electronic device, characterized in that: include: A memory storing an executable program; A processor, connected to the memory via a bus, and configured to run the program, wherein the program executes the method described in any one of claims 1 to 9 when running.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 9.

12. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 9.

Citation Information

Cited By

  • Method and program product for providing question and answer service based on video content

    CN121262422A