Method, system and device for controlling program playing of virtual host and medium

Through the control method of virtual hosts playing programs, combined with content analysis, background information generation, sentiment analysis and intelligent recommendation, the problem of inability to provide personalized services in the existing technology is solved, and personalized content recommendation and user experience improvement is achieved.

CN120343304APending Publication Date: 2025-07-18SHENZHEN COOCAA NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510191302.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing virtual host playback method cannot provide personalized services through simple speech recognition and text reading.

Method used

The control method of virtual host playing programs is adopted, including content analysis module, background information generation module, intelligent highlight recommendation module, user interaction module and emotion analysis module. By obtaining user input information, analyzing content data, generating background information, analyzing user emotions and interest modes, personalized playback content is comprehensively recommended.

Benefits of technology

It realizes the personalized content experience based on user needs and interests, and improves the user's viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120343304A_ABST
    Figure CN120343304A_ABST
Patent Text Reader

Abstract

The invention relates to a method, a system and equipment for controlling a virtual host to play a program, and a medium. The method comprises the following steps: acquiring input information of a user and determining an intention recognition result of the input information; using the content analysis module to obtain and analyze to-be-processed content data based on the input information to obtain an analysis result; generating background information of the to-be-processed content data based on a preset database by using the background information generation module; obtaining a personalized portrait of a user and an interested content mode, and inputting the personalized portrait, the content mode, the analysis result and the background information into the intelligent watching point recommendation module to obtain a recommendation result output by the intelligent watching point recommendation module; and obtaining the target playing content of the virtual host based on the intention recognition result and the recommendation result. According to the invention, personalized user experience can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technologies, and in particular, to a control method, system, device, and medium for a virtual host to play programs. Background Art

[0002] With the continuous progress of artificial intelligence technologies, virtual assistants and intelligent systems have been widely applied in multiple fields.

[0003] In the field of television program and movie content introduction, the related technology is that a virtual host plays program content for users based on simple speech recognition and text reading functions.

[0004] However, simple speech recognition and text reading cannot provide personalized services to users. Summary of the Invention

[0005] Embodiments of this application provide a control method, system, device, and medium for a virtual host to play programs, aiming to solve the technical problem that personalized services cannot be provided to users due to the virtual host playing method using simple speech recognition and text reading functions.

[0006] In a first aspect, embodiments of this application provide a control method for a virtual host to play programs. The control system for a virtual host to play programs includes a content parsing module, a background information generation module, an intelligent highlight recommendation module, and a user interaction module; the user interaction module is used to execute the following method steps, and the method includes:

[0007] Obtain the input information of the user and determine the intention recognition result of the input information;

[0008] Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result;

[0009] Use the background information generation module to generate background information of the content data to be processed based on a preset database;

[0010] Obtain the personalized portrait and the content pattern of interest of the user, input the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlight recommendation module, and obtain the recommendation result output by the intelligent highlight recommendation module;

[0011] Based on the intention recognition result and the recommendation result, obtain the target playing content of the virtual host.

[0012] A further technical solution thereof is that the control system for a virtual host to play programs further includes an emotion analysis module, and the method further includes:

[0013] Input the input information into the sentiment analysis module for sentiment analysis to obtain a sentiment analysis result;

[0014] The step of inputting the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlights recommendation module to obtain the recommendation result output by the intelligent highlights recommendation module includes:

[0015] Input the personalized portrait, the content pattern, the parsing result, the background information, and the sentiment analysis result into the intelligent highlights recommendation module to obtain the recommendation result output by the intelligent highlights recommendation module.

[0016] A further technical solution thereof is that the step of inputting the input information into the sentiment analysis module for sentiment analysis to obtain a sentiment analysis result includes:

[0017] Use the sentiment analysis module to perform voice sentiment analysis, text sentiment analysis, and expression sentiment analysis on the input information to obtain voice sentiment, text sentiment, and expression sentiment;

[0018] Integrate the voice sentiment, the text sentiment, and the expression sentiment to obtain a sentiment analysis result.

[0019] A further technical solution thereof is that the step of obtaining the user's personalized portrait and the content pattern of interest includes:

[0020] Obtain the user's viewing history data, click behavior data, and viewing evaluation data;

[0021] Generate a personalized portrait of the user based on the viewing history data, the click behavior data, and the viewing evaluation data;

[0022] Extract high-frequency highlight segments based on the viewing history data and the viewing evaluation data by using data mining technology;

[0023] Determine the content pattern of interest to the user from the high-frequency highlight segments.

[0024] A further technical solution thereof is that the step of using the background information generation module to generate the background information of the content data to be processed based on a preset database includes:

[0025] Obtain relevant information from a preset database, where the relevant information includes at least one of director background information, actor background information, screenwriter background information, relevant interview information, and behind-the-scenes footage information;

[0026] Preprocess the relevant information to obtain preprocessed background information;

[0027] Match the preprocessed background information with the content data to be processed to obtain a matching result;

[0028] If the matching result meets the preset requirements, use the preprocessed background information as the background information of the content data to be processed.

[0029] A further technical solution thereof is that the preprocessing of the relevant information to obtain preprocessed background information includes:

[0030] Clean, de-duplicate, and structure the relevant information to obtain preprocessed background information.

[0031] A further technical solution thereof is that the parsing result includes a text parsing result and a video parsing result, and obtaining the content data to be processed and parsing it to obtain a parsing result includes:

[0032] Perform natural language processing on the text data in the content data to be processed to obtain a text parsing result, where the text parsing result includes a plot summary;

[0033] Use computer vision technology to parse the video data in the content data to be processed and mark key events in the video data to obtain a video parsing result.

[0034] In a second aspect, an embodiment of the present application further provides a control system for a virtual host to play a program. The control system for a virtual host to play a program includes a content parsing module, a background information generation module, an intelligent highlight recommendation module, a user interaction module, and an emotion analysis module;

[0035] Among them, the content parsing module is used to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result;

[0036] The background information generation module is used to generate the background information of the content data to be processed based on a preset database;

[0037] The emotion analysis module is used to perform emotion analysis on the input information to obtain an emotion analysis result;

[0038] The intelligent highlight recommendation module is used to obtain a recommendation result based on the personalized portrait, the content pattern, the parsing result, the background information, and the emotion analysis result.

[0039] The user interaction module is used to obtain the input information of the user and determine the intention recognition result of the input information, and obtain the target playback content of the virtual host based on the intention recognition result and the recommendation result.

[0040] In a third aspect, an embodiment of the present application further provides a computer device, which includes a memory and a processor. A computer program is stored on the memory, and when the processor executes the computer program, the above method is implemented.

[0041] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium. The storage medium stores a computer program, and when the computer program is executed by a processor, the above method can be implemented.

[0042] An embodiment of the present application provides a control method, system, device and medium for a virtual host to play programs. Among them, the method includes: obtaining input information of a user and determining an intention recognition result of the input information; using the content parsing module to obtain content data to be processed based on the input information and perform parsing to obtain a parsing result; using the background information generation module to generate background information of the content data to be processed based on a preset database; obtaining a personalized portrait of the user and an interested content pattern, and inputting the personalized portrait, the content pattern, the parsing result and the background information into the intelligent highlight recommendation module to obtain a recommendation result output by the intelligent highlight recommendation module; and obtaining target playing content of the virtual host based on the intention recognition result and the recommendation result.

[0043] By determining the intention recognition result of the input information, an embodiment of the present application can understand the user's needs. By obtaining the personalized portrait of the user and the interested content pattern, it can deeply understand the user's interests. And by parsing the program content and providing detailed background information, the intelligent highlight recommendation module can perform comprehensive analysis according to the personalized portrait, the content pattern, the parsing result and the background information, recommend a recommendation result that conforms to the user's personalized content, and then obtain the target playing content of the virtual host according to the intention recognition result and the recommendation result. In this way, a personalized content experience can be provided for the user. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0045] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0046] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated. The drawings in the figures do not constitute a scale limitation.

[0047] Figure 1 Schematic flowchart of the first embodiment of a control method for a virtual host to play a program provided by the present application;

[0048] Figure 2 Interaction diagram of each module in the control system for a virtual host to play a program provided by the present application;

[0049] Figure 3 Schematic structural diagram of an embodiment of a computer device provided by the present application. Detailed implementation manners

[0050] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0051] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. This repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0052] It should be understood that when used in this specification and the appended claims, the terms "comprises" and "comprising" indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0053] It should also be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include the plural forms.

[0054] It should also be further understood that the term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0055] As used in this specification and the appended claims, the term "if" may be construed, depending on the context, as "when" or "once" or "in response to determining" or "in response to detecting". Similarly, the phrase "if determined" or "if [the described condition or event] is detected" may be construed, depending on the context, as meaning "once determined" or "in response to determining" or "once [the described condition or event] is detected" or "in response to detecting [the described condition or event]".

[0056] With the continuous progress of artificial intelligence technology, virtual assistants and intelligent systems have been widely applied in many fields.

[0057] In the specific field of television program and movie content introduction, the related technology is that a virtual host plays program content for users based on simple speech recognition and text reading functions.

[0058] However, simple speech recognition and text reading cannot provide personalized services to users.

[0059] To solve the technical problem in the prior art that the virtual host playback method using simple speech recognition and text reading functions cannot provide personalized services to users, the present application provides a control method for a virtual host to play programs, which can provide users with a personalized content experience.

[0060] Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of a control method for a virtual host to play programs provided by the present application. The method includes:

[0061] Step 110: Obtain the input information of the user and determine the intention recognition result of the input information.

[0062] Step 120: Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result.

[0063] Step 130: Use the background information generation module to generate the background information of the content data to be processed based on a preset database.

[0064] Step 140: Obtain the personalized portrait and the content pattern of interest of the user, and input the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module.

[0065] Step 150: Based on the intention recognition result and the recommendation result, obtain the target playback content of the virtual host.

[0066] In this way, by determining the intention recognition result of the input information, the user's needs can be understood. By obtaining the user's personalized portrait and the content patterns of interest, the user's interests can be deeply understood. And by parsing the program content and providing detailed background information, the intelligent highlights recommendation module can perform comprehensive analysis based on the personalized portrait, the content pattern, the parsing result, and the background information, recommend a recommendation result that conforms to the user's personalized content, and then obtain the target playback content of the virtual host according to the intention recognition result and the recommendation result. In this way, a personalized content experience can be provided to the user.

[0067] In some embodiments, the control system for the virtual host to play programs further includes an emotion analysis module. For details, refer to the second embodiment. The second embodiment includes steps 210 - step 260:

[0068] Step 210: Obtain the user's input information and determine the intention recognition result of the input information.

[0069] Among them, the input information may be "Please introduce the highlights of 'XXXX'.", and it can be determined that the intention recognition result of this input information is "The user wants to know the exciting scenes of the movie."

[0070] Step 220: Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result.

[0071] For example, if the content data to be processed is the script and video data of "XXXX", the parsing result may include key scenes, important actions, main characters, etc.

[0072] In some embodiments, the parsing result includes a text parsing result and a video parsing result. Obtaining the content data to be processed and performing parsing in step 220 to obtain a parsing result includes:

[0073] Step 221: Perform natural language processing on the text data in the content data to be processed to obtain a text parsing result, where the text parsing result includes a plot summary.

[0074] Step 222: Use computer vision technology to parse the video data in the content data to be processed and mark the key events in the video data to obtain a video parsing result.

[0075] Step 230: Use the background information generation module to generate the background information of the content data to be processed based on a preset database.

[0076] In some embodiments, in step 230, using the background information generation module to generate the background information of the content data to be processed based on a preset database includes steps 231 - 234:

[0077] Step 231: Obtain relevant information from the preset database, where the relevant information includes at least one of director background information, actor background information, screenwriter background information, relevant interview information, and behind - the - scenes footage information.

[0078] Step 232: Pre - process the relevant information to obtain the pre - processed background information.

[0079] In some embodiments, the relevant information can be cleaned, de - duplicated, and structured to obtain the pre - processed background information.

[0080] Step 233: Match the pre - processed background information with the content data to be processed to obtain a matching result.

[0081] Step 234: If the matching result meets the preset requirements, use the pre - processed background information as the background information of the content data to be processed.

[0082] Step 240: Input the input information into the sentiment analysis module for sentiment analysis to obtain a sentiment analysis result.

[0083] In some embodiments, step 240, that is, inputting the input information into the sentiment analysis module for sentiment analysis to obtain a sentiment analysis result, includes:

[0084] Step 241: Use the sentiment analysis module to perform voice sentiment analysis, text sentiment analysis, and facial expression sentiment analysis on the input information to obtain voice sentiment, text sentiment, and facial expression sentiment.

[0085] Step 242: Integrate the voice sentiment, the text sentiment, and the facial expression sentiment to obtain a sentiment analysis result.

[0086] Step 250: Obtain the user's personalized portrait and the content pattern of interest, and input the personalized portrait, the content pattern, the parsing result, the background information, and the sentiment analysis result into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module.

[0087] Step 260: Based on the intention recognition result and the recommendation result, obtain the target playback content of the virtual host.

[0088] In this way, by using the sentiment analysis module to perform sentiment analysis on the input information, a sentiment analysis result is obtained; furthermore, the intelligent highlight recommendation module performs comprehensive analysis based on the personalized portrait, the content pattern, the parsing result, the background information, and the sentiment analysis result, and recommends a recommendation result that conforms to the user's personalized content, thereby improving the accuracy of the content recommended by the intelligent highlight recommendation module.

[0089] In some embodiments, obtaining the user's personalized portrait and the content pattern of interest in step 250 may include the following steps:

[0090] Step 251: Obtain the user's viewing history data, click behavior data, and viewing evaluation data.

[0091] Step 252: Generate a personalized portrait of the user based on the viewing history data, the click behavior data, and the viewing evaluation data.

[0092] Step 253: Extract high-frequency highlight segments by using data mining techniques based on the viewing history data and the viewing evaluation data.

[0093] Step 254: Determine the content pattern of interest to the user from the high-frequency highlight segments.

[0094] In addition, the present application also provides a control system for a virtual host to play programs. The control system for a virtual host to play programs includes a content parsing module, a background information generation module, an intelligent highlight recommendation module, a user interaction module, and a sentiment analysis module.

[0095] Among them, the content parsing module is used to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result;

[0096] The background information generation module is used to generate background information of the content data to be processed based on a preset database;

[0097] The sentiment analysis module is used to perform sentiment analysis on the input information to obtain a sentiment analysis result;

[0098] The intelligent highlight recommendation module is used to obtain a recommendation result based on the personalized portrait, the content pattern, the parsing result, the background information, and the sentiment analysis result.

[0099] The user interaction module is used to obtain the input information of the user and determine the intention recognition result of the input information, and obtain the target playback content of the virtual host based on the intention recognition result and the recommendation result.

[0100] Specifically, the following introductions are made to each of the above modules:

[0101] 1. The content analysis module is used to extract valuable information from text and video data for subsequent processing and recommendation, including text analysis and video analysis specifically.

[0102] Among them, text analysis includes the following:

[0103] 1) Data acquisition: Obtain the script and subtitle data of programs, movies or TV dramas from content providers through the API.

[0104] 2) Natural Language Processing (NLP): Includes word segmentation, Named Entity Recognition (NER) and relation extraction. Word segmentation is used to split text content into words or phrases, NER is used to identify entities such as people, places, events, etc., and relation extraction is used to identify the relationships between entities (such as the association between the protagonist and the supporting role, and the main events).

[0105] 3) Plot summary generation: Generate a plot summary using the extracted key information to facilitate users to quickly understand the content outline.

[0106] Video analysis includes the following:

[0107] 1) Video data acquisition: Obtain video files through the API.

[0108] 2) Computer Vision (CV) technology: Includes scene recognition, action recognition and expression recognition. Scene recognition is used to mark important scene transitions, action recognition is used to identify key actions (such as combat scenes, hugs, etc.), and expression recognition is used to analyze the emotions of characters.

[0109] 3) Event marking: Mark the identified important scenes and actions to ensure the comprehensiveness of content analysis.

[0110] 2. The background information generation module is used to retrieve through the Internet and internal databases to generate background information related to the current program, mainly including Internet data scraping and internal database retrieval.

[0111] Among them, Internet data scraping includes:

[0112] 1) Retrieval mechanism: The system uses web crawler technology to scrape background information of directors, actors, screenwriters, etc. from the Internet, as well as relevant interviews, behind-the-scenes footage and other information.

[0113] 2) Data processing: Clean, deduplicate and structure the scraped data to ensure the accuracy and readability of the data.

[0114] Internal database retrieval includes:

[0115] 1) Database: Includes historical databases (such as information on past programs and people), encyclopedic knowledge bases, etc.

[0116] 2) Dynamic update: Regularly update the database content to ensure the provision of the latest background information.

[0117] 3) Semantic matching: Utilize semantic analysis technology to match the retrieved background information with the current content to ensure relevance.

[0118] 3. The intelligent highlight recommendation module is used to provide personalized highlight recommendations by combining big data analysis and user behavior analysis, mainly including big data analysis and user behavior analysis.

[0119] Among them, big data analysis includes:

[0120] 1) Data collection: Collect the viewing data and evaluation data of a large number of users.

[0121] 2) Data mining: Use data mining technology to extract high-frequency highlights and bright spots, such as the high evaluation times of a certain plot and the popular discussion degree of a certain character.

[0122] 3) Pattern recognition: Use machine learning algorithms to identify the content patterns that users are commonly interested in.

[0123] User behavior analysis includes:

[0124] 1) Viewing history: Record the user's viewing history, including the type of programs watched, duration, frequency, etc.

[0125] 2) Click behavior: Analyze the user's click behavior on the interface to identify the user's preferences.

[0126] 3) Feedback collection: Collect the user's feedback information, such as ratings, comments, etc.

[0127] 4) Personalized portrait: Generate a user's personalized portrait based on the user's historical data and behavior analysis for precise recommendation.

[0128] 4. The user interaction module is used to achieve diverse interactions with users by using speech recognition and natural language understanding technologies, mainly including speech recognition, natural language understanding, and natural language generation.

[0129] Among them, speech recognition includes:

[0130] 1) Real-time recognition: Use high-precision speech recognition technology to convert the user's speech input into text.

[0131] 2) Speech model training: Train the model based on a large-scale speech dataset to improve the recognition accuracy and response speed.

[0132] Natural language understanding (NLU) includes:

[0133] 1) Intent recognition: Analyze the user's speech text to identify the user's intent (such as asking about the plot, background information, highlight recommendations, etc.).

[0134] 2) Dialogue management: Manage the dialogue process based on the user's intent to ensure the coherence and naturalness of the interaction.

[0135] Natural language generation (NLG) includes:

[0136] 1) Answer generation: Use NLG technology to generate appropriate answer text.

[0137] 2) Speech synthesis: Convert the generated text into natural speech through speech synthesis technology for broadcasting.

[0138] 5. The sentiment analysis module is used to identify the user's emotional state by analyzing the user's speech, text, and expressions to optimize the interaction experience, mainly including speech sentiment analysis, text sentiment analysis, and expression sentiment analysis.

[0139] Among them, speech sentiment analysis includes:

[0140] 1) Speech feature extraction: Extract emotional features in the speech, such as pitch, speech rate, volume, etc.

[0141] 2) Sentiment classification: Use the sentiment classification model to judge the user's emotional state (such as happy, sad, angry, etc.).

[0142] Text sentiment analysis includes:

[0143] 1) Sentiment dictionary: Construct a sentiment dictionary containing positive, negative, and neutral sentiment words.

[0144] 2) Text analysis: Analyze the emotional tendency of the user's text input, and make an emotional judgment by combining the sentiment dictionary and context information.

[0145] Expression sentiment analysis includes:

[0146] 1) Expression recognition: Collect the user's expression data through the camera, and use the expression recognition technology to judge the user's emotional state.

[0147] 2) Emotional feedback: Combine the results of speech and text sentiment analysis to comprehensively judge the user's overall emotional state.

[0148] Based on the introduction of the above various modules, combined with Figure 2 , the control method for the virtual host to play the program provided by this application mainly includes the following:

[0149] 1) The user interaction module obtains the user's input information from the user side. Among them, the input information can be camera capture, user voice activation, or text information directly input by the user.

[0150] Among them, when the input information is voice information, the user interaction module performs speech recognition on the collected input information to convert the input information into text information.

[0151] 2) Use natural language understanding technology to identify and understand the text information to obtain the intention recognition result.

[0152] 3) The content parsing module obtains script data and video data from the content library data, inputs the script data into the text parsing sub-module for natural language processing to obtain the text parsing result output by the text parsing sub-module, and the text parsing result includes the plot summary; the video data has been input into the video parsing sub-module, and computer vision technology is used to parse the video data, and key events and key actions in the video data are marked to obtain the final video parsing result; and the video parsing result and the text parsing result are respectively input into the intelligent recommendation module.

[0153] 4) The background information generation module grabs relevant background data from the Internet and performs preprocessing to obtain the preprocessed background data.

[0154] The internal database retrieves according to the historical knowledge base (such as previous program and character information), the encyclopedia knowledge base, etc. to obtain the retrieved background information, matches the retrieved background information with the current content, and provides the background information that meets the matching requirements to the intelligent highlight recommendation module.

[0155] Among them, the internal database receives the relevant background data grabbed from the Internet and preprocessed to regularly update the database content to ensure the provision of the latest background information.

[0156] 5) The emotion analysis module inputs the user input information into the speech emotion analysis sub-module, the text emotion analysis sub-module, and the expression emotion analysis sub-module respectively for analysis to obtain the speech emotion analysis result output by the speech emotion analysis sub-module; the text emotion analysis result output by the text emotion analysis sub-module, and the expression emotion analysis result output by the expression emotion analysis sub-module, and inputs the speech emotion analysis result, the text emotion analysis result, and the expression emotion analysis result into the intelligent emotion analysis integration module, so that the intelligent emotion analysis integration module outputs the emotion analysis result to the intelligent recommendation module.

[0157] 6) The intelligent recommendation module generates user personalized tags by obtaining viewing history data, click behavior data, and feedback collection data, and then obtains the user's personalized portrait, and mines and performs pattern recognition on the collected data to obtain the content patterns that the user is interested in;

[0158] And based on the personalized portrait, the content pattern, the parsing result, the background information, and the sentiment analysis result, an intelligent personalized recommendation result is output to the user interaction module.

[0159] 7) The user interaction module performs speech synthesis based on the intelligent personalized recommendation result and the intention recognition result to obtain the content finally played by the virtual host, that is, the target playback content.

[0160] Exemplarily, assuming that the virtual host introduces the movie "XXXX", a control method for a virtual host to play a program provided by the present application may include the following processes:

[0161] 1) The content parsing module obtains the script and video data of "XXXX" through the API interface. Using NLP technology, the plot summary, main characters, and important events of the movie are extracted. At the same time, computer vision technology identifies key scenes and important actions.

[0162] 2) The background information generation module grabs and generates background information related to "XXXX", including the introduction of the director (XXX), the backgrounds of the main actors, the filming story of the movie, the adaptation information of the original comic, etc.

[0163] 3) The intelligent highlight recommendation module analyzes big data and finds that users have extremely high evaluations of highlights such as the exciting scenes of XXX. At the same time, user behavior analysis shows that the current user prefers action scenes, so the intelligent highlight recommendation module focuses on recommending these highlights of "XXX".

[0164] 4) The user interaction module receives the user's voice query: "Please introduce the highlights of "XXXX".", identifies the user's intention and parses that the user wants to know the exciting scenes of the movie. The virtual host answers in an oral way: "The final battle in "XXXX" is very exciting, especially the sacrifice of XX which shocked countless audiences. There is also the duel between X and XX, which is definitely not to be missed!" At the same time, the tone of the host is adjusted according to the user's emotional state to make the answer more vivid and contagious.

[0165] 5) The sentiment analysis module detects that the user shows some dull emotions when hearing the sacrifice of XX, so the virtual host adds: "Although the departure of XX makes people feel heartache, his heroic fearlessness saved the universe and will always live in our hearts!"

[0166] Based on the above, the present application realizes the intelligent, humanized, and personalized content introduction and recommendation functions through the content parsing module, the background information generation module, the intelligent highlight recommendation module, the user interaction module, and the sentiment analysis module, thereby improving the user's movie-watching experience.

[0167] Such as Figure 3As shown in the figure, an embodiment of the present application provides a computer device, including a processor 111, a communication interface 112, a memory 113, and a communication bus 114. Among them, the processor 111, the communication interface 112, and the memory 113 complete mutual communication through the communication bus 114.

[0168] The memory 113 is used to store computer programs.

[0169] In an embodiment of the present application, when the processor 111 is used to execute the program stored on the memory 113, it implements the control method for a virtual host to play a program provided by any of the foregoing method embodiments, including:

[0170] Obtain the input information of the user and determine the intention recognition result of the input information;

[0171] Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result;

[0172] Use the background information generation module to generate the background information of the content data to be processed based on a preset database;

[0173] Obtain the user's personalized portrait and the content pattern of interest, and input the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module;

[0174] Based on the intention recognition result and the recommendation result, obtain the target playback content of the virtual host.

[0175] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a storage medium, and this storage medium is a computer-readable storage medium. This computer program is executed by at least one processor in the computer system to implement the process steps of the method embodiments of the above method.

[0176] Therefore, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the control method for a virtual host to play a program provided by any of the foregoing method embodiments, including:

[0177] Obtain the input information of the user and determine the intention recognition result of the input information;

[0178] Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result;

[0179] Generate the background information of the content data to be processed by the background information generation module based on a preset database;

[0180] Obtain the personalized portrait of the user and the content patterns of interest, and input the personalized portrait, the content patterns, the parsing result, and the background information into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module;

[0181] Based on the intention recognition result and the recommendation result, obtain the target playback content of the virtual host.

[0182] The storage medium is a physical, non-transitory storage medium. For example, it can be various physical storage media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disc that can store program codes. The computer-readable storage medium can be non-volatile or volatile.

[0183] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0184] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0185] The steps in the method embodiments of this application can be adjusted, combined, and deleted according to actual needs. The units in the device embodiments of this application can be combined, divided, and deleted according to actual needs. In addition, the functional units in each embodiment of this application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.

[0186] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0187] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0188] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, provided that these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application also intends to include these changes and modifications.

[0189] As described above, the above are only the specific implementation manners of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of various equivalent modifications or substitutions, and these modifications or substitutions should all be covered by the protection scope of this application. Therefore, the protection scope of this application shall be subject to the protection scope of the claims.

Claims

1. A control method for a virtual host to play programs, characterized in that, The control system for a virtual host to play programs includes a content parsing module, a background information generation module, an intelligent highlight recommendation module, and a user interaction module; The user interaction module is used to execute the following method steps, and the method includes: Obtain the input information of the user and determine the intention recognition result of the input information; Use the content parsing module to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result; Use the background information generation module to generate the background information of the content data to be processed based on a preset database; Obtain the user's personalized portrait and the content pattern of interest, and input the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module; Based on the intention recognition result and the recommendation result, obtain the target playback content of the virtual host.

2. The method according to claim 1, wherein The control system for a virtual host to play programs further includes an emotion analysis module, and the method further includes: Input the input information into the emotion analysis module for emotion analysis to obtain an emotion analysis result; The step of inputting the personalized portrait, the content pattern, the parsing result, and the background information into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module includes: Input the personalized portrait, the content pattern, the parsing result, the background information, and the emotion analysis result into the intelligent highlight recommendation module to obtain the recommendation result output by the intelligent highlight recommendation module.

3. The method according to claim 1, wherein The step of inputting the input information into the emotion analysis module for emotion analysis to obtain an emotion analysis result includes: Use the emotion analysis module to perform voice emotion analysis, text emotion analysis, and expression emotion analysis on the input information to obtain voice emotion, text emotion, and expression emotion; Integrate the voice emotion, the text emotion, and the expression emotion to obtain an emotion analysis result.

4. The method according to claim 1, wherein The step of obtaining the user's personalized portrait and the content pattern of interest includes: Obtain the user's viewing history data, click behavior data, and viewing evaluation data; Generate the user's personalized portrait based on the viewing history data, the click behavior data, and the viewing evaluation data; Extract high-frequency highlight segments based on the viewing history data and the viewing evaluation data using data mining techniques; Determine the content pattern of interest to the user from the high-frequency highlight segments.

5. The method according to claim 1, wherein The step of using the background information generation module to generate the background information of the content data to be processed based on a preset database includes: Obtain relevant information from a preset database, and the relevant information includes at least one of director background information, actor background information, screenwriter background information, relevant interview information, and behind-the-scenes footage information; Perform preprocessing on the relevant information to obtain preprocessed background information; Match the preprocessed background information with the content data to be processed to obtain a matching result; If the matching result meets the preset requirements, use the preprocessed background information as the background information of the content data to be processed.

6. The method according to claim 5, wherein Preprocessing the relevant information to obtain preprocessed background information, including: Cleaning, de-duplicating, and structuring the relevant information to obtain preprocessed background information.

7. The method according to claim 1, wherein The parsing result includes a text parsing result and a video parsing result. Obtaining the content data to be processed and performing parsing to obtain the parsing result, including: Performing natural language processing on the text data in the content data to be processed to obtain a text parsing result, where the text parsing result includes a plot summary. Using computer vision technology to parse the video data in the content data to be processed and marking key events in the video data to obtain a video parsing result.

8. A control system for a virtual host to play programs, characterized in that, The control system for the virtual host to play a program includes a content parsing module, a background information generation module, an intelligent highlight recommendation module, a user interaction module, and an emotion analysis module. Among them, the content parsing module is used to obtain the content data to be processed based on the input information and perform parsing to obtain a parsing result. The background information generation module is used to generate background information for the content data to be processed based on a preset database. The emotion analysis module is used to perform emotion analysis on the input information to obtain an emotion analysis result. The intelligent highlight recommendation module is used to obtain a recommendation result based on the personalized portrait, the content pattern, the parsing result, the background information, and the emotion analysis result. The user interaction module is used to obtain the input information of the user and determine the intention recognition result of the input information, and obtain the target playing content of the virtual host based on the intention recognition result and the recommendation result.

9. A computer device, characterized in that, The computer device includes a memory and a processor. A computer program is stored on the memory. When the processor executes the computer program, the method described in any one of claims 1-7 is implemented.

10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program. When the computer program is executed by the processor, the method described in any one of claims 1-7 can be implemented.