Intelligent interaction method and system in film watching scene, electronic equipment and storage medium
By generating an AI agent based on film footage and user interaction records in a movie-watching scenario, the problem of low interaction efficiency in existing technologies is solved, achieving a highly efficient user interaction experience and enhanced immersion.
Patent Information
- Application Number
- CN202511391849.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-02-17
AI Technical Summary
In existing technologies, existing AI agents suffer from low interaction efficiency in complex task processing and user interaction. The loss of context caused by switching between existing AI agents necessitates users repeatedly providing the same background information to different agents, resulting in reduced interaction efficiency and a poor user experience.
In a movie-watching scenario, the system responds to user interaction commands, determines the target movie based on the currently displayed screen content, acquires related data and the user's historical interaction records, identifies the target character based on the screen content, related data, and the user's historical interaction records, generates an AI agent corresponding to the target character, receives user input interaction content, performs semantic recognition on the interaction content, invokes the AI agent, generates response content based on the semantic recognition results, and displays the response content in combination with the target character's role characteristics.
It improves the efficiency of interaction in movie-watching scenarios, enhances the user's immersion and emotional connection, and improves the user's interactive experience by generating a highly realistic AI agent to converse with the user.
Smart Images

Figure CN121541772A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of artificial intelligence, and in particular to an intelligent interaction method and system in a viewing scenario, an electronic device, and a storage medium. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, AI agent systems have been widely applied in customer service, knowledge question answering, personalized recommendation, and other fields. Existing AI agent technology mainly responds to user queries through pre-trained models (such as large language models LLM) or rule-based systems.
[0003] Although the current technology can support context tracking for a single dialogue (for example, a chat robot maintains short-term conversation memory), context loss occurs due to agent switching in complex task processing, and users need to repeatedly provide the same background information to different agents, resulting in reduced interaction efficiency and poor user experience. SUMMARY
[0004] Embodiments of the present application provide an intelligent interaction method and system in a viewing scenario, an electronic device, and a storage medium to at least solve the problem of reduced interaction efficiency in related technologies.
[0005] In a first aspect, embodiments of the present application provide an intelligent interaction method in a viewing scenario, the method comprising: In response to a user interaction instruction, determining a target film according to the current display content of the device, obtaining associated data of the target film and historical interaction records of the user; Determining a target role according to the display content, the associated data, and the historical interaction records, and generating an AI agent corresponding to the target role; Receiving interaction content input by the user, performing semantic recognition on the interaction content, calling the AI agent, generating response content according to the semantic recognition result, and displaying the response content in combination with the role characteristics of the target role.
[0006] In some embodiments, determining a target role according to the display content, the associated data, and the historical interaction records comprises: Analyzing the display content to obtain a character role contained in the display content; Analyzing the associated data to obtain character roles in the target film and importance of each character role; Determining a love value of each character role of the user according to the historical interaction records; Determining a confidence of each character role based on the character role contained in the display content, the importance, and the love value, and taking the character role with the highest confidence as the target role.
[0007] In some embodiments, the generating the AI agent corresponding to the target character comprises: obtaining the image features, the background story and the voice feature library of the target character from the association data; generating the AI agent based on the image features, the background story and the voice feature library.
[0008] In some embodiments, the displaying the response content in combination with the character features of the target character comprises: generating a text dialogue box at a preset position of the display interface, and displaying the response content in the text dialogue box; and / or calling a speech synthesis engine to output a voice reply with emotional intonation according to the character features of the target character and the response content.
[0009] In some embodiments, the method further comprises: generating a semi-transparent interactive layer in the display interface, wherein a virtual image of the target character is embedded in the semi-transparent interactive layer, and the virtual image supports triggering dynamic expression changes according to the response content; when the voice reply with emotional intonation is output, driving the mouth movement of the virtual image to match the voice waveform data in real time, and presenting a speaking state.
[0010] In some embodiments, the method further comprises: switching the target character to a user-specified character in response to a character switching instruction of the user; generating an AI agent based on the user-specified character for intelligent interaction.
[0011] In some embodiments, the method further comprises: periodically obtaining updated movies within a preset time period from a local movie database, and keywords of the movies, wherein the keywords include movie names and character names; based on the keywords, grabbing related content from multiple movie data sources to obtain association data of the movies.
[0012] In a second aspect, the embodiments of the present application provide an intelligent interaction system in a movie watching scenario, the system comprising: a data acquisition module configured to determine a target movie according to current displayed content of a picture in response to a user interaction instruction, and acquire association data of the target movie and historical interaction records of the user; an agent generation module configured to determine a target character based on the picture content, the association data and the historical interaction records, and generate an AI agent corresponding to the target character. The interaction module is configured to receive the interaction content input by the user, perform semantic recognition on the interaction content, call the AI agent, generate response content according to the semantic recognition result, and display the response content in combination with the role characteristics of the target role.
[0013] In a third aspect, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the intelligent interaction method in a viewing scenario as described in the first aspect when executing the computer program.
[0014] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program executable by a processor to implement the intelligent interaction method in a viewing scenario as described in the first aspect.
[0015] Compared with the related art, the intelligent interaction method in a viewing scenario provided by the embodiment of the present application responds to the user interaction instruction, determines the target film according to the current display content of the device, obtains the associated data of the target film and the historical interaction record of the user, determines the target role according to the picture content, the associated data and the historical interaction record, and generates the AI agent corresponding to the target role, receives the interaction content input by the user, performs semantic recognition on the interaction content, calls the AI agent, generates the response content according to the semantic recognition result, and displays the response content in combination with the role characteristics of the target role, thereby solving the problem of reduced interaction efficiency in a viewing scenario. In the user viewing state, the AI agent can be aroused for chatting. The AI agent customizes the role AI agent according to the role information, the film plot information and the historical interaction record of the user, the customized AI agent combines the personality characteristics and the dialogue style of the role, and carries out dialogue with the user, thereby improving the interaction efficiency and enhancing the user interaction experience. BRIEF DESCRIPTION OF DRAWINGS
[0016] The accompanying drawings, which are included to provide a further understanding of the present application, constitute a part of the present application, and the illustrative embodiments of the present application and their description serve to explain the present application, and do not constitute improper limitations on the present application. In the drawings: Figure 1 is a flowchart of the intelligent interaction method in a viewing scenario according to an embodiment of the present application; Figure 2 is a dialog box schematic diagram according to an embodiment of the present application; Figure 3 is a structural block diagram of the intelligent interaction system in a viewing scenario according to an embodiment of the present application; Figure 4 is an internal structure schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to make the purposes, technical solutions, and advantages of the present application clearer, the present application is described and explained below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application. Based on the embodiments provided by the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort belong to the scope of the present application.
[0018] It is obvious that the accompanying drawings in the following description are only some examples or embodiments of the present application, and for those of ordinary skill in the art, the present application can be applied to other similar scenarios without creative effort based on the accompanying drawings. In addition, it can be understood that although the efforts made in the development process can be complex and lengthy, for those of ordinary skill in the art related to the content disclosed in the present application, some designs, manufacturing or production changes based on the technical content disclosed in the present application are only routine technical means and should not be understood as insufficient disclosure of the content disclosed in the present application.
[0019] In the present application, "embodiments" means that the specific features, structures or properties described in conjunction with the embodiments can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive or alternative embodiments to other embodiments. It is explicitly and implicitly understood by those of ordinary skill in the art that the embodiments described in the present application can be combined with other embodiments without conflict.
[0020] Unless otherwise defined, technical terms and scientific terms used in the present application shall have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. Unless otherwise defined, the terms "one", "a", "an", "the" and like terms refer to and encompass both singular and plural referents. The terms "comprising", "having", "including", and "containing" and their variations, are intended to be open-ended terms that encompass both the stated specification and additional elements not specified. For example, a process, method, system, product, or apparatus that comprises a list of steps or elements is not limited to only those steps or elements but can include additional steps or elements not expressly listed, or can also include additional steps or elements that are inherent in the process, method, system, product, or apparatus. The terms "connected" and "coupled" and their variations are intended to encompass a connection between two or more devices, which can be either a direct connection or an indirect connection via one or more additional devices and connections. The term "multiple" means two or more. The term "and / or" describes associated objects, which can exist in three states: single, combined, and separate. The character " / " generally represents an "or" relationship between the associated objects. The terms "first", "second", "third", and the like are used only to distinguish similar objects, and do not represent a specific order.
[0021] The present embodiment provides an intelligent interaction method in a movie watching scenario. Figure 1 The flowchart of the intelligent interaction method in a movie watching scenario according to the present embodiment is shown in FIG. 1, which includes the following steps: Figure 1 The flowchart of the intelligent interaction method in a movie watching scenario according to the present embodiment is shown in FIG. 1, which includes the following steps: Step S101, in response to a user interaction instruction, determining a target movie according to the current display content of the device, obtaining associated data of the target movie and historical interaction records of the user.
[0022] The user issues an interaction instruction during the movie watching process, obtains relevant data of the currently played movie and historical interaction records of the user, and generates an AI agent to interact with the user.
[0023] Step S102, determining a target character according to the picture content, the associated data and the historical interaction records, and generating an AI agent corresponding to the target character.
[0024] In some embodiments, the step S102 of determining a target character according to the picture content, the associated data and the historical interaction records includes: Step S1021, analyzing the picture content to obtain the character in the picture content.
[0025] Step S1022, analyzing the associated data to obtain the characters in the target movie and the importance of each character.
[0026] At step S1023, the user's favorite value for each character role is determined according to the historical interaction record.
[0027] At step S1024, the confidence of each character role is determined based on the character roles contained in the picture content, the importance and the favorite value, and the character role with the highest confidence is taken as the target role.
[0028] In this embodiment, the target role can only be selected from the characters in the current picture content, or all characters in the film can be selected as the selection object.
[0029] For example, the frame analysis module is called to intercept the current video frame, and the pre-trained visual recognition model (optionally, YOLOv7+ResNet model) is used to detect the main body of the picture: characters A and B are recognized.
[0030] In the case of only selecting characters in the current picture content as the selection object: the importance of characters A and B is determined according to the associated data, and the user's favorite value for characters A and B is determined according to the user's historical interaction record. Optionally, in the form of weighted sum, the confidence of characters A and B is obtained based on the importance and the favorite value. For example, the confidence of character A is 98%, and the confidence of character B is 95%, then character A is taken as the target role. Wherein, the importance of the character can be determined according to the proportion of the character's appearance time in the film, the correlation with the main plot and the character's popularity (such as the access amount of the character's entry in each platform).
[0031] Only the confidence of the character in the picture needs to be analyzed, which is efficient and fast in response. By weighting and integrating the "importance of the character" and the "user's favorite value", both the narrative logic of the content itself (the importance of the character in the current scene) and the user's personal emotion are respected.
[0032] In the case of all characters in the film as the selection object: in the case where characters A and B are not the main characters of the film, the main characters of the film are determined: characters C and D. Characters A, B, C and D are all taken as the alternative of the target role, and each character is attached with an initial weight value. The initial weight value of the character in the current picture (characters A and B) is higher than that of the character not in the picture, and the confidence of each character is determined according to the initial weight value of the character, the importance and the user's favorite value of the character.
[0033] Higher initial weight value is given to the characters in the picture to ensure that the recommendation list still focuses on the current picture; the final confidence is determined by the initial weight value, the global importance and the user's favorite value. This makes the result not only close to the current context, but also takes into account the role in the whole work and the user's personal historical preference.
[0034] In some embodiments, the step S102 of generating an AI agent corresponding to the target role includes: Step S1025, obtaining the image characteristics, background story and voice feature library of the target role from the associated data.
[0035] Step S1026, generating an AI agent based on the image characteristics, background story and voice feature library.
[0036] Query the role database to obtain the 3D image data, background story, and voice feature library (e.g., tone is a teenage voice line, and speech speed is 1.2 times) of the target role. Load the knowledge graph submodule to obtain associated data, which includes: structured relationship data (e.g., the relationship between role A and role B), film script dialogue library (e.g., role A's address to role B), and encyclopedia knowledge entries. Based on the obtained data, generate an AI agent instance and assign a unique ID (e.g., A_Agent_20231025-1430).
[0037] Step S103, receiving user input interaction content, performing semantic recognition on the interaction content, calling an AI agent, generating response content based on the semantic recognition result, and displaying the response content in combination with the role characteristics of the target role.
[0038] For example, the user asks through voice: "What is the relationship between role A and role B?" The system performs: converts the user's semantic question into text through speech recognition (ASR), extracts key entities: [role A, role B, character relationship] through a semantic understanding model (optionally, BERT + domain entity recognition); Activate the A_Agent_20231025-1430 instance to generate a natural language answer that conforms to the personality of the target role from the knowledge graph search.
[0039] In some embodiments, the step S103 of displaying the response content in combination with the role characteristics of the target role includes: Step S1031, generating a text dialogue box at a preset position on the display interface and displaying the response content in the text dialogue box.
[0040] Step S1032, calling a speech synthesis engine to output a voice reply with emotional intonation based on the role characteristics of the target role and the response content.
[0041] Multi-modal output control: superimpose a semi-transparent interactive layer on the projection interface, display a Q version of character A with dynamic expressions. At the same time, display a text dialogue box on the display page to output the response content in the form of text. Optionally, the text dialogue box is positioned at the lower right corner of the picture, the font is Kai, and the font size is adaptive. Synchronously call the speech synthesis engine to output a voice reply with emotional intonation through Tacotron2 and the voiceprint characteristics of character A. In the process of voice reply, the Q version of the image reflects the speaking state. Figure 2 is a dialogue box schematic diagram according to an embodiment of the present application.
[0042] In some embodiments, the method further comprises: Generating a semi-transparent interactive layer on the display interface, embedding a virtual character image of the target character in the semi-transparent interactive layer, and supporting dynamic expression changes triggered according to the response content.
[0043] When outputting a voice reply with emotional intonation, driving the mouth movement of the virtual character image to match the voice waveform data in real time, and presenting the speaking state.
[0044] The semi-transparent effect avoids rough cutting of the UI elements on the picture, giving a light and high-tech feeling that the information layer "floats" above the original content, and is more comfortable and high-level visually. Triggering dynamic expressions (such as smiling, surprised, thinking) according to the response content is the key to conveying emotions and understanding, which goes beyond a simple "talking machine", making the character appear alive and emotional, and enhancing emotional resonance.
[0045] The interactive layer can display the text record of the dialogue for the user to read or review carefully; the voice reply with emotional intonation directly conveys emotions and highlights; the expression and lip shape of the virtual image provide visual assistance and reinforcement for voice information. This "sound, text, and shape" integrated interaction method conforms to the natural communication habits of humans, and has very high information transmission efficiency. Users do not need to exert effort to understand, but can naturally absorb information, greatly reducing the cognitive burden of interaction.
[0046] Through the above steps, in the user viewing state, the AI intelligent body chat can be aroused, the AI intelligent body is customized according to the character information, the film plot information, and the user historical interaction record, the customized AI intelligent body combines the personality characteristics and the dialogue style of the character, and carries on the dialogue with the user, improves the interaction efficiency, and improves the user interaction experience. The generated AI intelligent body is no longer a pile of labels, but a highly realistic, consistent, and credible "digital life body", which enhances the user's sense of immersion and emotional connection.
[0047] In some embodiments, the method further comprises: In response to the character switching instruction of the user, switching the target character to the user-specified character; Generate AI agent based on user specified role for intelligent interaction.
[0048] The user may not be interested in the target role recommended by the system at the moment, therefore, the user can also actively switch the role corresponding to the AI agent according to his / her current idea in this embodiment. The system is responsible for reducing the selection cost of the user (intelligent recommendation), but does not deprive the user of the final selection right, realizing the balance between "intelligent recommendation" and "user control".
[0049] In some embodiments, the method further comprises: Periodically obtaining updated movies in a preset time period from the local movie database, and keywords of the movies, the keywords including movie names and role names.
[0050] Based on the keywords, related content is grabbed from multiple movie and television data sources to obtain associated data of the movie.
[0051] The system can periodically (such as every week) scan the updates of the local database to find movies that will be released or have just been released, and pre-grab trailers, news and known role information of these movies for preprocessing and preloading. When the user issues an interaction signal, the interaction service can be provided in the first time.
[0052] It should be noted that the steps shown in the above process or the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from here.
[0053] The embodiment also provides an intelligent interaction system in a movie watching scenario, which is used to implement the above embodiments and preferred embodiments, and will not be described again. As used below, the terms "module", "unit", "sub-unit" and the like can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware or a combination of software and hardware is also possible and is conceived.
[0054] Figure 3 is a structural block diagram of the intelligent interaction system in a movie watching scenario according to the embodiment of the application, as Figure 3 shown, the system comprises: The data acquisition module 31 is configured to determine a target movie according to the current display content of the device in response to a user interaction instruction, and acquire associated data of the target movie and a historical interaction record of the user.
[0055] The agent generation module 32 is configured to determine a target role according to the display content, the associated data and the historical interaction record, and generate an AI agent corresponding to the target role.
[0056] The interaction module 33 is configured to receive the interaction content input by the user, perform semantic recognition on the interaction content, invoke the AI agent, generate the response content according to the semantic recognition result, and display the response content in combination with the role characteristics of the target role.
[0057] In some embodiments, the agent generation module 32 includes: A picture recognition module is configured to analyze the picture content to obtain the character roles included in the picture content.
[0058] A role analysis module is configured to analyze the association data to obtain the character roles in the target movie and the importance of each character role.
[0059] A like value determination module is configured to determine the like value of each character role of the user according to the historical interaction record.
[0060] A target role determination module is configured to determine the confidence of each character role based on the character roles included in the picture content, the importance, and the like value, and determine the character role with the highest confidence as the target role.
[0061] In some embodiments, the agent generation module 32 includes: A data acquisition module is configured to acquire the image characteristics, the background story, and the voice feature library of the target role from the association data.
[0062] An agent construction module is configured to generate the AI agent based on the image characteristics, the background story, and the voice feature library.
[0063] In some embodiments, the interaction module 33 includes: A visual interaction module is configured to generate a text dialogue box at a preset position of the display interface and display the response content in the text dialogue box.
[0064] A voice interaction module is configured to invoke a voice synthesis engine, output a voice reply with emotional intonation according to the role characteristics of the target role and the response content.
[0065] In some embodiments, the visual interaction module further includes: A first visual interaction module is configured to generate a semi-transparent interaction layer on the display interface, embed a virtual character image of the target role in the semi-transparent interaction layer, and support dynamic expression changes of the virtual character image according to the response content.
[0066] A second visual interaction module is configured to drive the mouth movement of the virtual character image to be real-time matched with the voice waveform data when the voice reply with emotional intonation is output, and present a speaking state.
[0067] In some embodiments, the system further comprises a role switching module configured to switch the target role to a user-specified role in response to a user role switching instruction, and generate an AI agent based on the user-specified role for intelligent interaction.
[0068] In some embodiments, the system further comprises a data updating module configured to periodically obtain updated movies in a preset time period and keywords of the movies, including movie names and role names, from the local movie database, and based on the keywords, obtain associated data of the movies by scraping related content from multiple movie and television data sources.
[0069] Through the above system, the AI agent can be aroused for chatting in the user's movie watching state. The AI agent is customized based on role information, movie plot information, and user historical interaction records. The customized AI agent combines the personality characteristics and dialogue style of the role to have a dialogue with the user, thereby improving the interaction efficiency and enhancing the user interaction experience. The generated AI agent is no longer a pile of labels, but a highly realistic, consistent, and credible "digital life form", which enhances the user's immersion and emotional connection.
[0070] It should be noted that each of the above modules can be a functional module or a program module, which can be implemented by software or hardware. For modules implemented by hardware, each of the above modules can be located in the same processor; or each of the above modules can also be located in different processors in any combination.
[0071] The embodiment also provides an electronic device including a memory and a processor, the memory storing a computer program, and the processor being configured to execute the computer program to perform the steps in any of the above method embodiments.
[0072] Optionally, the electronic device can further include a transmission device and an input and output device, wherein the transmission device is connected to the processor, and the input and output device is connected to the processor.
[0073] Optionally, in the embodiment, the processor can be configured to execute the following steps through the computer program: S1, in response to a user interaction instruction, determining a target movie based on the current display content of the device, and obtaining associated data of the target movie and historical interaction records of the user.
[0074] S2, determining a target role based on the picture content, the associated data and the historical interaction records, and generating an AI agent corresponding to the target role.
[0075] S3, receiving the interactive content input by the user, performing semantic recognition on the interactive content, calling an AI agent, generating response content according to the semantic recognition result, and displaying the response content in combination with the role characteristics of the target role.
[0076] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementation manners, and this embodiment will not be repeated here.
[0077] In one embodiment, Figure 4 is a schematic diagram of the internal structure of an electronic device according to an embodiment of the application, as Figure 4 indicated, an electronic device is provided, which can be a server, and the internal structure diagram of the electronic device can be as Figure 4 indicated. The electronic device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the electronic device is used to provide computing and control capability. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the electronic device is used to store data. The network interface of the electronic device is used to communicate with the external terminal through the network connection. The computer program is executed by the processor to implement an intelligent interaction method in a viewing scene.
[0078] Those skilled in the art can understand, Figure 4 the structure shown in the figure, only the block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device can include more or less components than those shown in the figure, or combine certain components, or have a different component arrangement.
[0079] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. Any reference to memory, storage, database or other medium used in each embodiment provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0080] Those skilled in the art should understand that each technical feature of the above-mentioned embodiments can be combined arbitrarily, and in order to make the description simple, each technical feature in the above-mentioned embodiments is not described all possible combinations, however, as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present application.
[0081] The above-mentioned embodiments only express several implementation manners of the present application, and the description is more specific and detailed, but it should not be understood as a limitation on the scope of the patent. It should be pointed out that for ordinary skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the patent of the present application should be subject to the appended claims.
Claims
1. An intelligent interaction method in a viewing scene, characterized in that, The method comprises: in response to a user interaction instruction, determining a target film according to the current display content of the device, obtaining the associated data of the target film and the historical interaction record of the user; determining a target role according to the picture content, the associated data and the historical interaction record, and generating an AI agent corresponding to the target role; receiving user input interaction content, performing semantic recognition on the interaction content, calling the AI agent, generating response content according to the semantic recognition result, and combining the role characteristics of the target role to display the response content.
2. The method of claim 1, wherein, The determination of the target role according to the picture content, the associated data and the historical interaction record comprises: analyzing the picture content to obtain the character roles contained in the picture content; analyzing the associated data to obtain the character roles in the target film and the importance of each character role; determining the user's love value for each character role according to the historical interaction record; based on the character roles contained in the picture content, the importance and the love value, determining the confidence of each character role, and taking the character role with the highest confidence as the target role.
3. The method of claim 1, wherein, The generation of the AI agent corresponding to the target role comprises: obtaining the image characteristics, background story and voice feature library of the target role from the associated data; based on the image characteristics, the background story and the voice feature library, generating the AI agent.
4. The method of claim 1, wherein, The display of the response content in combination with the role characteristics of the target role comprises: generating a text dialogue box at a preset position of the display interface, and displaying the response content in the text dialogue box; and / or calling a speech synthesis engine to output a voice reply with emotional intonation according to the role characteristics of the target role and the response content.
5. The method of claim 2, wherein, The method further comprises: generating a semi-transparent interaction layer in the display interface, embedding a virtual character image of the target role in the semi-transparent interaction layer, and the virtual character image supports triggering dynamic expression changes according to the response content; when the voice reply with emotional intonation is output, driving the mouth movement of the virtual character image to match the voice waveform data in real time, and presenting a speaking state.
6. The method of claim 1, wherein, The method further comprises: in response to a user role switching instruction, switching the target role to a user specified role; generating an AI agent based on the user specified role for intelligent interaction.
7. The method of claim 1, wherein, The method further comprises: periodically obtaining updated films in a preset time period and keywords of the films from a local film database, the keywords including film names and role names; based on the keywords, grabbing related content from multiple film and television data sources to obtain the associated data of the films.
8. An intelligent interaction system in a viewing scene, characterized in that, The system comprises: a data acquisition module configured to determine a target film according to the current display content of the device in response to a user interaction instruction, and obtain the associated data of the target film and the historical interaction record of the user; an agent generation module configured to determine a target role according to the picture content, the associated data and the historical interaction record, and generate an AI agent corresponding to the target role; An interaction module is configured to receive an interaction content input by a user, perform semantic recognition on the interaction content, invoke the AI agent, generate a response content according to a result of the semantic recognition, and display the response content in combination with a role feature of the target role.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the intelligent interaction method in the viewing scenario as claimed in any one of claims 1 to 7.
10. A storage medium having stored thereon a computer program, characterized in that The program is executed by the processor to implement the intelligent interaction method in the viewing scenario as claimed in any one of claims 1 to 7.
Citation Information
Patent Citations
Information interaction method and device, electronic equipment and storage medium
CN116680376A
Virtual image display method and related equipment
CN117119236A
Display device, server and interactive processing method
CN119440374A
The Method Of Providing A Character Call, Computing System For Performing The Same, And Computer-Readable Recording Medium
KR102738804B1
KR20250128632A