Cross-device collaborative film watching interaction method and device, smart television and storage medium
By collaborating with smart TVs and audio output devices, user voice data is acquired and matched to generate interactive audio, solving the problem of lacking personalized narration in existing smart TV viewing modes, improving user experience and maintaining a sense of immersion in the viewing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN ZHIXIAN VISION SOFTWARE TECHNOLOGY CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-01
AI Technical Summary
The existing smart TV viewing modes lack personalized narration, resulting in a poor user experience. Furthermore, interactive functions are prone to conflict with the audio of movies and TV shows, affecting the immersive viewing experience.
By working in conjunction with smart TVs and audio output devices, preset plot elements of the media data selected by the user are obtained, the user's voice data is received and matched to generate response text data, and then converted into interactive audio data and distributed to the audio output device for playback, ensuring separation from film and television audio and avoiding conflicts.
It enables personalized narration without interfering with the audio source of the main viewing device, enhancing the viewing experience and maintaining the user's immersion in the movie.
Smart Images

Figure CN121967802A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent device control technology, and in particular to a cross-device collaborative viewing interaction method, device, smart TV and storage medium. Background Technology
[0002] With the widespread adoption of artificial intelligence technology, in order to enhance product competitiveness, large-scale models can now be used to enrich the playback functions of smart TVs and improve the user's viewing experience.
[0003] However, existing smart TV viewing modes are mostly one-way playback, lacking personalized interaction and emotional companionship with users; even if interactive functions exist, they are mostly produced by the TV alone, which can easily conflict with the audio of movies and TV shows, causing auditory discomfort.
[0004] Therefore, the existing movie-watching mode cannot achieve personalized narration without interfering with the audio source of the main movie-watching device, resulting in a poor user experience. Summary of the Invention
[0005] The main purpose of this application is to provide a cross-device collaborative viewing interaction method, device, smart TV and storage medium, which aims to solve the technical problem that existing viewing modes cannot achieve personalized narration without interfering with the audio source of the viewing device.
[0006] To achieve the above objectives, this application proposes a cross-device collaborative viewing interaction method, which is applied to a smart TV connected to an audio output device; the method includes: Obtain the preset plot elements corresponding to the media data selected by the user to be played; The system receives user voice data collected by the audio output device, performs data matching between the user voice data and the preset plot elements, and generates response text data based on the matching results. The response text data is converted into interactive audio data, and the interactive audio data is distributed to the audio output device for playback when the media data to be played is output.
[0007] Furthermore, to achieve the above objectives, this application also proposes a cross-device collaborative viewing interaction device, which is applied to a smart TV connected to an audio output device. The device includes: The data management module is used to obtain the preset plot elements corresponding to the media data selected by the user to be played; The device configuration module is used to select at least one audio output device from the audio output devices in the current environment; The response processing module is used to receive user voice data collected by the audio output device, perform data matching between the user voice data and the preset plot elements, and generate response text data based on the matching results. The data distribution module is used to convert the response text data into interactive audio data, and distribute the interactive audio data to the audio output device for playback when outputting the media data to be played.
[0008] In addition, to achieve the above objectives, this application also proposes a smart TV, which includes: a memory, a processor, and a cross-device collaborative viewing interaction program stored on the memory and executable on the processor. The cross-device collaborative viewing interaction program is configured to implement the steps of the cross-device collaborative viewing interaction method described above.
[0009] In addition, to achieve the above objectives, this application also provides a storage medium storing a program for implementing a cross-device collaborative viewing interaction method, wherein the program for implementing the cross-device collaborative viewing interaction method is executed by a processor to implement the steps of the cross-device collaborative viewing interaction method as described above.
[0010] This application provides a cross-device collaborative viewing interaction method, device, smart TV, and storage medium. The method includes: acquiring preset plot elements corresponding to media data selected by the user; receiving user voice data collected by an audio output device, performing data matching between the user voice data and the preset plot elements, and generating response text data based on the matching result; converting the response text data into interactive audio data, and distributing the interactive audio data to the audio output device for playback when outputting the media data to be played.
[0011] This application proposes an innovative movie-watching companion solution by integrating with the existing smart hardware ecosystem. This application utilizes an auxiliary device, i.e., an audio output device, independent of the main movie-watching device (i.e., the smart TV's main audio source), to play interactive voice based on pre-processed plot information while the user is watching the media to be played. This achieves multimodal voice interaction while maintaining the user's immersion in the movie, providing a superior viewing experience. Attached Figure Description
[0012] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart illustrating the first embodiment of the cross-device collaborative movie-watching interaction method of this application; Figure 2 This is a schematic diagram illustrating the mode switching process of the cross-device collaborative movie-watching interaction method of this application; Figure 3 This is a schematic diagram of the companion mode interaction process of the cross-device collaborative movie-watching interaction method of this application; Figure 4 This is a flowchart illustrating the second embodiment of the cross-device collaborative movie-watching interaction method of this application; Figure 5 This is a schematic diagram illustrating the model invocation process of the cross-device collaborative movie-watching interaction method of this application; Figure 6 This is a schematic diagram of the module structure of the cross-device collaborative movie-watching interaction device according to an embodiment of this application; Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the cross-device collaborative movie-watching interaction method in this application embodiment.
[0015] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0016] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0017] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0018] The main solution of this application is to propose a cross-device collaborative viewing interaction method for smart TVs. The smart TV is connected to an audio output device, and the smart TV obtains preset plot elements corresponding to the media data to be played selected by the user; it receives user voice data collected by the audio output device, performs data matching between the user voice data and the preset plot elements, and generates response text data based on the matching result; it converts the response text data into interactive audio data, and distributes the interactive audio data to the audio output device for playback when outputting the media data to be played.
[0019] In existing movie-watching modes, when users require narration, the main audio source (the film's original soundtrack) on the smart TV and the accompanying voice can easily mix, causing auditory interference and reducing the immersive viewing experience. Therefore, existing movie-watching modes lack cross-device collaborative interaction capabilities and cannot achieve personalized narration without interfering with the main movie-watching device's audio source.
[0020] To address this issue, this application proposes an innovative movie-watching companion solution that integrates with the existing smart hardware ecosystem. This application utilizes an auxiliary device—an audio output device—independent of the primary movie-watching device (i.e., the smart TV's main audio source), to play interactive voice based on pre-processed plot information while the user is watching the media to be played. This achieves multimodal voice interaction while maintaining the user's immersion in the movie, providing a superior viewing experience.
[0021] It should be noted that the executing entity in this embodiment can be a cross-device collaborative viewing interaction system, or a television terminal with network connectivity, voice recognition, model deployment, and device communication functions, capable of loading and playing media data, and able to interact with audio output devices and cloud servers, such as tablet computers, personal computers, mobile phones, and smart TVs, or a cross-device collaborative viewing interaction device capable of achieving the above functions. This embodiment does not specifically limit it. The following description uses a smart TV as the executing entity as an example to illustrate this embodiment and the following embodiments, where the smart TV is connected to an audio output device.
[0022] It is easy to understand that the aforementioned cross-device collaboration can represent that a smart TV can establish communication connections with multiple audio output devices (such as smartphones, tablets, Bluetooth headsets, smart speakers, etc.) and achieve collaborative voice output through data transmission and command scheduling. The core of this process lies in the smart TV acting as the control center and the audio output devices acting as interactive audio playback carriers. The audio output devices can be electronic devices with voice capture (such as mobile phones with microphones or smart speakers) and audio playback functions, determined by the smart device through a list of supported output devices in the surrounding environment. They do not require independent internet access to the cloud and only serve as voice input and audio output carriers.
[0023] Based on this, embodiments of this application provide a cross-device collaborative movie-watching interaction method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the cross-device collaborative movie-watching interaction method of this application.
[0024] In this embodiment, the cross-device collaborative viewing interaction method includes steps S10~S30: Step S10: Obtain the preset plot elements corresponding to the media data to be played selected by the user; It is easy to understand that the aforementioned media data to be played can be user-selected movies, TV series, or other film and television content, stored in the smart TV's local film and television library or a cloud-based film and television database. The preset plot elements can be a structured data set pre-organized for the media data to be played, covering core elements such as time, characters, plot, and environment. In this embodiment, the preset plot elements may include a plot outline and a complete script, and are pre-stored in the smart TV's local film and television library as accompanying data of the media data to be played.
[0025] It should be understood that in this embodiment, the smart TV can be configured with a "normal mode" and a "companion mode," and the smart TV will only provide subsequent voice interaction and companionship based on preset plot elements when the user selects the "companion mode." For ease of understanding, refer to... Figure 2 The data acquisition process in this embodiment will be explained and described. Figure 2 This diagram illustrates the mode switching process of the cross-device collaborative viewing interaction method of this application. Figure 2 As shown, users can select a movie or story as the media to be played using the smart TV's remote control, touchscreen, or voice commands. While the smart TV loads this media data, it can simultaneously access the local film and television library or the cloud-based film and television database to retrieve the preset plot elements corresponding to the media data to be played. Understandably, Figure 2 The "mode permission" refers to whether the source of the media data to be played meets the conditions for "companion mode". The brief content is the "plot outline" of the media data to be played. It will be downloaded to the smart TV along with the film and television resources and used to perform information matching when querying the conditions for "companion mode" on the device side.
[0026] At this point, the smart TV can determine whether the media data to be played has a "companion mode". For example, if the media data does not have preset plot elements, it indicates that "companion mode" is not supported, and the smart TV directly enters "normal mode". If it has preset plot elements, the smart TV can further prompt the user that "companion mode is supported" (such as a screen pop-up window + voice prompt). The prompt can use a countdown method (such as a 10-second countdown). If the user still does not take any action after the countdown ends, i.e., does not switch, it defaults to "normal mode". At the same time, during the subsequent playback of content with "companion mode", if the user actively chooses to switch, the smart TV can still enter "companion mode" to continue the subsequent process.
[0027] Step S20: Receive user voice data collected by audio output device, perform data matching between user voice data and preset plot elements, and generate response text data based on the matching results; Understandably, the aforementioned user voice data can be natural language speech information collected by the user through the microphone of an audio output device, which may include questions, emotional expressions, and interactive requests from the user while watching the media data to be played. In this case, the aforementioned data matching can refer to the process by which the smart TV uses a retrieval and recall algorithm to calculate the semantic similarity between the user voice data and preset plot elements in order to determine the degree of association between the two.
[0028] For example, after entering companion mode, the smart TV can send a "voice capture ready" command to the connected audio output device. The audio output device then activates its microphone to capture the user's voice data in real time and transmits it to the smart TV. The smart TV uses its built-in ASR (Automatic Speech Recognition) module to convert the user's voice data into text information, and then calls a retrieval algorithm to semantically match the text information with preset plot elements, calculate the semantic similarity value (value range 0~1), and determine whether the response conditions are met.
[0029] It is easy to understand that in the current playback environment, there may be multiple audio output devices that can communicate with a smart TV. This not only makes it difficult for users to quickly locate the playback device for interactive voice communication, but also makes it difficult to clearly identify the interactive terminal, which can lead to chaotic voice acquisition and hinder subsequent real-time data interaction processing. Therefore, in a feasible implementation, in this embodiment, step S20 may include steps A1 to A5: Step A1: Broadcast a device discovery request to the audio output device and generate a list of available devices corresponding to the audio output device in response to the device discovery request; the list of available devices includes paired devices and unpaired devices; Understandably, the aforementioned device discovery request can be a broadcast signal sent by a smart TV to search for nearby connectable audio output devices, supporting common communication protocols such as Bluetooth and Wi-Fi. In this case, the list of available devices can be a device list generated by the smart TV based on the response to the device discovery request, and may include nearby connectable audio output devices. Furthermore, in this embodiment, the list of available devices can be displayed in categories of "paired devices" and "unpaired devices," with paired devices appearing at the top.
[0030] For example, such as Figure 2As shown, after a smart TV enters "Companion Mode," it can automatically send a device discovery request (such as Bluetooth broadcast + Wi-Fi LAN broadcast) to search for nearby audio output devices that support the corresponding communication protocol. Upon receiving the request, nearby audio output devices, if in a connectable state, will return a response signal (including device name, device type, communication address, etc.). The smart TV collects all response signals, generates a list of available devices that supports multiple selections, and displays the device names on the screen. It can also mark devices as "paired" or "unpaired," and sort paired devices by connection history (most recently connected devices are listed first).
[0031] Step A2: Determine the selected target paired device and / or target unpaired device based on the user's selection instruction for the list of available devices; Understandably, the aforementioned selection command can be an operation command issued by the user via smart TV remote control, touch screen, or voice, selecting the target playback device from the list of available devices. Therefore, the target paired device can be an audio output device that the user has selected from the list of available devices and has previously been successfully paired with the smart TV; while the target unpaired device can be an audio output device that the user has selected from the list of available devices and has not been paired with the smart TV.
[0032] Step A3: If there is an unpaired target device, verify the pairing identity of the unpaired target device and select the unpaired target device that passes the pairing identity verification as the interactive playback terminal. Step A4: If a target paired device exists, select the target paired device as the interactive playback terminal; Step A5: Receive user voice data collected by the interactive playback terminal.
[0033] It should be noted that the aforementioned pairing identity verification refers to the verification process for establishing a connection between the smart TV and the target unpaired device. In this embodiment, the smart TV can verify the pairing identity of the user-selected target unpaired device through methods such as Bluetooth pairing code, Wi-Fi password verification, or QR code verification to ensure the security of the device connection. If the verification passes, the target unpaired device is selected as the interactive playback terminal; if the verification fails (e.g., due to an incorrect pairing code), the smart TV displays "Pairing failed" and allows the user to reselect the device or retry pairing. Furthermore, the user-selected target paired device does not require a redundant identity verification process, simplifying the operation.
[0034] Therefore, the aforementioned interactive playback terminal can be an audio output device that, after user selection and identity verification, is determined to be used for collecting user voice data and playing interactive audio data. Figure 2 The accompanying voice output device shown. For example... Figure 2As shown, the output device list can be determined based on the user's selection command for the available device list. If the output device list only contains the target paired device, the smart TV will directly select it as the interactive playback terminal without additional identity verification and control the target paired device to issue a confirmation prompt tone. If the output device list only contains the target unpaired device, the user can proceed to the pairing identity verification as instructed, and select the target unpaired device that passes the pairing identity verification as the interactive playback terminal. The user can also select the corresponding role language style type for the target unpaired device. If both the target paired device and the target unpaired device that passes the verification are selected at the same time, both will be determined as interactive playback terminals.
[0035] In this embodiment, the smart TV can automatically broadcast device discovery requests, eliminating the need for users to manually enter search mode and effectively simplifying the connection process. Simultaneously, the list of available devices can be sorted by "paired / unpaired," allowing users to quickly locate target devices, improving selection efficiency. Furthermore, verifying the pairing identity of unpaired devices ensures secure device connections, preventing unauthorized access from unfamiliar devices. Therefore, this embodiment effectively identifies the interactive playback terminal, ensuring the accuracy and stability of voice capture and providing reliable voice input support for subsequent interactions.
[0036] Step S30: Convert the response text data into interactive audio data, and distribute the interactive audio data to the audio output device for playback when outputting the media data to be played.
[0037] It is understandable that the aforementioned response text data can be natural language text generated by the smart TV based on data matching results, which fits the plot development and user needs, and may include plot explanations, emotional responses, and interesting explanations.
[0038] For example, after determining the semantic matching result, if the similarity value reaches a preset threshold (e.g., 0.15), it indicates that the user's voice is strongly related to the plot. The smart TV can generate response text data based on the corresponding information fragments in the preset plot elements and the user's voice intent (e.g., if the user asks "Why did character 1 do this?", an explanation text is generated based on character 1's behavioral motivation in the plot elements). If the similarity value does not reach the preset threshold, it can further determine whether it is necessary to call the cloud server to supplement information before generating response text data.
[0039] At this point, the interactive audio data can be a playable audio signal generated by the smart TV after converting the response text data through TTS (Text To Speech) technology, which can include control information such as character voice, tone, and rhythm.
[0040] Therefore, during the distribution of interactive audio data, the smart TV can act as a control hub, sending the interactive audio data to the selected audio output device according to preset rules to ensure that interactive audio playback and media data output do not conflict. For example, the smart TV can convert response text data into interactive audio data through a TTS module. During the conversion process, the tone, speed, and emotional tone of the voice can be adjusted according to the user's preset role style (such as lively, calm, or humorous). Simultaneously, the smart TV can control the audio output of the media data to be played (maintaining the original volume without reducing or interfering with it) and distribute the interactive audio data to the selected audio output device, which then independently plays the interactive audio, achieving a separate playback mode where "video audio is output from the TV, and interactive audio is output from a portable device."
[0041] It is understandable that existing interactive audio conversion lacks personalized design and suffers from a monotonous voice style. To address this issue, in a feasible implementation, step S30 in this embodiment may include steps B1-B3: Step B1: Convert the response text data into audio source text data according to the user-selected role control parameters, and obtain the target role voice data packet from the preset role voice database based on the audio source text data; It is important to understand that the aforementioned role control parameters can be parameters selected by the user to define the interactive audio style, and may include role type (such as friend, narrator, character in the play), tone (lively, calm, funny, empathetic), speaking speed (fast, medium, slow), pitch (high, medium, low), etc.
[0042] At this point, the aforementioned audio source text data can be optimized text generated by the smart TV based on the response text data and combined with role control parameters for TTS conversion. In this embodiment, the audio source text data may include the response text content and control information such as emotion, rhythm, and pauses required for speech synthesis (such as marking "pause 0.5 seconds" or "increased tone").
[0043] For example, if the generated response text data is "The plot twist was unexpected", then the audio source text data can be generated by combining the user-selected character control parameters ("lively girl", "medium speaking speed" and "empathic tone") as: "[empathy], [medium speaking speed], text 'This plot twist was so unexpected' and [pause 0.3 seconds]".
[0044] Understandably, the aforementioned preset character voice database can be a collection of voice data packets stored locally on the smart TV or in the cloud, and each voice data packet can correspond to a character style (such as "lively girl," "calm uncle," or "humorous commentator"), containing voice synthesis parameters for that character style. If the user does not select character control parameters, the "general friendly" character voice data packet can be used by default. In this case, the target character voice data packet can be a voice data packet selected from the preset character voice database that matches the character control parameters selected by the user.
[0045] Step B2: Generate interactive audio data based on the target character's voice data packet and response text data; It should be understood that in this embodiment, the smart TV can input the audio source text data into the pre-configured TTS module and load the target character's voice data packet at the same time, so that the TTS module can generate interactive audio data (such as an audio signal that conforms to the style of a "lively girl") based on the content and control information in the audio source text and the voice parameters of the character's voice data packet.
[0046] Step B3: Based on the device role identifier corresponding to the audio output device, when outputting media data to be played, distribute the interactive audio data to the audio output device for playback according to the playback timing strategy.
[0047] Understandably, the aforementioned device role identifier can be a unique identifier assigned by the smart TV to each selected audio output device, which can be used to distinguish the playback roles of different devices (such as device A corresponding to the "lively role" and device B corresponding to the "calm role").
[0048] The aforementioned playback timing strategy can be a set of rules for regulating the timing of interactive audio playback by multiple audio output devices to avoid audio conflicts. For example, in this embodiment, the playback timing strategy may include: a parallel playback strategy, where multiple devices simultaneously play the same or complementary interactive audio (such as a chorus-style response), suitable for scenarios that enhance the interactive atmosphere; or a sequential playback strategy, where multiple devices play interactive audio in a preset order (such as device A responding first, followed by device B responding later), suitable for multi-role conversational interactive scenarios.
[0049] For ease of understanding, please refer to Figure 3 The audio data distribution process in this embodiment will be explained and described. Figure 3 This is a schematic diagram of the companion mode interaction process of the cross-device collaborative viewing interaction method of this application. In this embodiment, the smart TV can determine the playback timing strategy based on the device role identifier corresponding to each audio output device. At this time, if the user selects only one audio output device, a single playback strategy can be directly adopted to distribute the interactive audio data to that device; if... Figure 3 As shown, the user selects multiple audio output devices (i.e. Figure 3 The mobile phone and portable speaker shown can select parallel or serial playback strategies according to the interactive scenario.
[0050] For example, a parallel playback strategy can be used during dramatic climaxes, allowing multiple audio playback devices to simultaneously respond to create tension; if the user selects a multi-character dialogue scene, a serial playback strategy can be used, distributing audio data sequentially according to the device's character identifier order. That is... Figure 3 As shown, if a smart TV is in companion mode and playing video content, and the user inputs "Is he in reality or a dream?" via far-field voice, the smart TV can receive the user's voice, understand the content, generate a response, and distribute the voices of different characters in an orderly manner. At this time, character A corresponding to the phone can respond: "I think he's in a dream because..."; character B can respond: "No, he's in reality, there's something on the screen...". Simultaneously, during the distribution of interactive audio data, the smart TV can ensure that the playback of interactive audio does not conflict with the audio output of the media data being played. For example, the interactive audio can be played during breaks in the video audio, or at a volume lower than the video audio without interfering with clear listening.
[0051] In this embodiment, by combining the role control parameters and the target character's voice data package, the interactive audio has a personalized style that matches user preferences; the playback timing strategy can standardize the playback logic of multiple devices, avoid audio conflicts, and improve the orderliness of the interaction; finally, in this embodiment, the interactive audio and the film and television audio are played separately and the playback timing is coordinated, which not only ensures the interactive effect but also does not interfere with the immersive viewing experience, thus improving the overall viewing experience.
[0052] In summary, addressing the issue that existing smart TV viewing modes are mostly one-way playback, lacking personalized interaction and emotional companionship with users, this embodiment can achieve personalized interaction that fits the plot by matching preset plot elements with user voice, simulating the empathetic response of human movie-watching communication, and enhancing the emotional viewing experience; at the same time, this embodiment can adopt a cross-device collaborative mode to separate the interactive audio from the film and television audio for playback, avoiding auditory interference and ensuring a sense of immersion in the viewing experience.
[0053] This embodiment provides a cross-device collaborative viewing interaction method applied to a smart TV, which is connected to an audio output device. The method includes: acquiring preset plot elements corresponding to media data selected by the user; broadcasting a device discovery request to the audio output device and generating a list of available devices corresponding to the audio output device in response to the device discovery request; the list of available devices includes paired devices and unpaired devices; determining the selected target paired device and / or target unpaired device according to the user's selection instruction on the list of available devices; if there is a target unpaired device, performing pairing identity verification on the target unpaired device and selecting the target unpaired device that passes the pairing identity verification as the interactive playback terminal; if there is a target paired device, selecting the target paired device as the interactive playback terminal; and receiving user voice data collected by the interactive playback terminal. The system performs data matching between user voice data and preset plot elements, and generates response text data based on the matching results. It then converts the response text data into audio source text data according to the user-selected role control parameters, and retrieves the target character's voice data package from the preset character voice database based on the audio source text data. Interactive audio data is generated based on the target character's voice data package and the response text data. Based on the device role identifier corresponding to the audio output device, the interactive audio data is distributed to the audio output device for playback according to the playback sequence strategy when outputting media data to be played. In this embodiment, addressing the issue that existing smart TV viewing modes are mostly one-way playback, lacking personalized interaction and emotional companionship with the user, this embodiment achieves personalized interaction that fits the plot by matching preset plot elements with the user's voice, simulating the empathetic response of human viewing communication, and enhancing the emotional viewing experience. Simultaneously, this embodiment can adopt a cross-device collaborative mode to separate the interactive audio from the film and television audio for playback, avoiding auditory interference and ensuring a sense of immersion in the viewing experience.
[0054] Based on the first embodiment of this application, in the second embodiment of this application, the same or similar content as the first embodiment described above can be referred to the above description, and will not be repeated hereafter.
[0055] It is easy to understand that in the current movie-watching interaction, the large model calls lack specificity. Regardless of whether the question is simple or complex, the cloud model is called, which can easily lead to high response delays and heavy load on the cloud server; or it can rely solely on the on-device accompanying language model, which cannot provide effective answers to complex questions due to insufficient information, thus affecting the user experience.
[0056] Therefore, based on the first embodiment, please refer to Figure 4 , Figure 4 This is a flowchart illustrating the second embodiment of the cross-device collaborative viewing interaction method of this application. In this embodiment, step S20 further includes steps C1 to C3: Step C1: Convert the user's voice data into query text, and perform semantic matching between the query text and preset plot elements to obtain semantic similarity values; It is easy to understand that after a smart TV receives user voice data transmitted from an audio output device, it can activate the built-in ASR module to convert the user voice data into natural language text, namely the query text mentioned above. This query text can completely preserve the user's core intent, such as questions and emotional expressions.
[0057] In the aforementioned semantic matching process, smart TVs can use NLP (Natural Language Processing) technology and retrieval recall algorithms to calculate the semantic similarity between the query text and preset plot elements. The query text is compared sentence-by-sentence with the outline and script content of the preset plot elements to calculate a semantic similarity value. This semantic similarity value quantifies the degree of association between the query text and the preset plot elements, ranging from 0 to 1, with higher values indicating stronger association.
[0058] Step C2: Determine the target response strategy based on the numerical comparison results between the preset similarity threshold and the semantic similarity value. The target response strategy includes the end-side companion language model invocation strategy or the cloud model invocation strategy. Step C3: Generate response text data based on the target response strategy and query text.
[0059] Understandably, the aforementioned preset similarity threshold can be a critical value used by the smart TV to determine whether it needs to call the cloud server. The specific value of this preset similarity threshold is preset by technical personnel (e.g., 0.15). In this case, if the semantic similarity value is higher than this threshold, it means that the information of the pre-built preset plot elements is sufficient to respond to the user's voice data; otherwise, it indicates that the information of the preset plot elements is insufficient to fully respond to the user's needs (e.g., the user asks about the historical background behind the plot or information about the actors), and the smart TV needs to supplement the cloud information to respond.
[0060] Therefore, in this embodiment, the target response strategy can be a response method determined based on the comparison result of semantic similarity value and preset similarity threshold. In this embodiment, the target response strategy may include a terminal-side accompanying language model invocation strategy (corresponding to the case when the semantic similarity value meets the standard) and a cloud-based model invocation strategy (corresponding to the case when the semantic similarity value does not meet the standard). That is, in this embodiment, different language models can be invoked to generate response text data corresponding to the user's voice data according to the actual situation.
[0061] Among them, the edge-side companion language model invocation strategy can be to directly call the language model deployed locally on the smart TV to generate response text data. Its advantage is that it does not need to access the cloud and has a fast response speed. The cloud model invocation strategy can be to upload relevant data to the cloud server and call the cloud companion language model to generate response text data. Its advantage is that it can obtain richer information support.
[0062] At this point, as one possible implementation method, in this embodiment, step C1 includes steps D1~D3: Step D1: When the target response strategy is the end-side companion language model invocation strategy, at least one information fragment with a semantic similarity value greater than the preset similarity threshold is obtained from the preset plot elements. Understandably, the aforementioned information fragments can be portions of the predefined plot elements whose semantic similarity to the query text exceeds a predefined similarity threshold, and may include specific fragments such as plot descriptions, character relationships, and scene backgrounds. For ease of understanding, please refer to... Figure 5 Provide an explanation. Figure 5 This is a schematic diagram illustrating the model invocation process of the cross-device collaborative viewing interaction method of this application. For example... Figure 5 As shown, after matching the user's dialogue content with the brief content of the source material to determine that the brief content of the source material meets the information content requirements for the response, and thus determining the target response strategy as the on-device companion language model invocation strategy, the smart TV can, based on the semantic similarity value obtained from the previous matching degree calculation, filter out all information fragments with semantic similarity values greater than a preset similarity threshold (such as 0.15) from the preset plot elements. For example, if the query text is "Why did character number 1 run away in the climax scene of a suspense film?", the smart TV can extract plot fragments related to "climax of a suspense film", "character number 1", and "escape motive" from the preset plot elements.
[0063] Step D2: Perform sentiment semantic analysis on the query text to obtain text prompt words; It is necessary to understand that, such as Figure 5 As shown, in the process of sentiment semantic analysis, smart TVs can rely on NLP technology (i.e., Figure 5 The natural language processing and language model application prompt word module shown performs emotion recognition on the query text and extracts the emotion type corresponding to the user's voice data (such as question, tension, joy, complaint, etc.). At this point, the aforementioned text prompt words can be guiding text generated by combining the core intent, emotion type, and information fragments of the query text. The on-device companion language model can be a lightweight large language model (LLM) deployed locally on the smart TV, capable of text generation, running without an internet connection, and with a fast response time.
[0064] For example, a smart TV can call its built-in emotional semantic analysis module to perform emotion recognition on the query text. If the user's emotion is identified as "tension," then it can generate text prompts containing empathetic guidance based on the information fragments, such as "The user is currently in a tense mood. Based on the background of character number 1 facing danger in the plot, an empathetic response is generated to explain the reason for their escape, with a tone that fits the tense atmosphere." If the user's emotion is "complaint," then the prompts can guide the generation of humorous explanatory text.
[0065] Step D3: Input the text prompts and information fragments into the locally deployed edge companion language model to obtain response text data.
[0066] It is easy to understand, such as Figure 5 As shown, the smart TV can encapsulate the generated text prompts and the selected target information fragments together as parameters for the model, and then input them into the on-device companion language model deployed locally on the smart TV. The on-device companion language model can generate response text data based on the prompts and the plot content of the information fragments. For example, for the above query text and prompts, the response text data generated is: "Aren't you super nervous? Character 1 has discovered that a villain is chasing him from behind, and the intersection ahead is the only escape route, so he must run quickly!"
[0067] In this embodiment, the smart TV can extract target information fragments to ensure that the response text generated by the on-device companion language model closely follows the plot and avoids deviating from the film and television content. At the same time, through emotional semantic analysis, the response text is made to match the user's emotions, achieving empathetic interaction. At this time, the interaction method based on the on-device companion language model has a fast response speed and no network latency, which can effectively improve the immediacy and user experience of the interaction.
[0068] It should be understood that the preset plot element information stored on the device side is limited and cannot effectively handle complex user questions regarding plot details, background knowledge, and extended interpretations. This results in complex queries failing to obtain effective responses, affecting the integrity of the interaction. Therefore, as an alternative implementation method, in this embodiment, the smart TV is connected to a cloud server, and step C1 includes steps E1~E2: Step E1: When the target response strategy is the cloud model invocation strategy, the media data to be played and the query text are uploaded to the cloud server so that the cloud server can generate optimized related data with expanded plot information based on the media data to be played and the query text, and input the optimized related data and the query text into the cloud companion language model to generate the response result. Step E2: Use the response result sent by the cloud server as the response text data.
[0069] It is easy to understand that the aforementioned cloud server can be a remote server with large-scale data storage, retrieval, model training, and text generation capabilities. In this embodiment, the cloud server can be connected to a smart TV via a network and stores a film and television database containing rich full-text film and television plot data, as well as network information data. At the same time, a cloud-based companion language model is deployed on the cloud server.
[0070] At this point, if the target response strategy is a cloud-based model invocation strategy, the smart TV can call the cloud backend access interface to upload the unique identifier of the media data to be played (such as the movie ID) and the user's query text to the cloud server via the network. The movie ID can be used by the cloud server to quickly locate the full plot data corresponding to the media data to be played, and the query text can be used to clarify the user's core needs.
[0071] It is important to understand that the aforementioned optimized related data can be a data set generated by the cloud server based on the media data to be played and the query text. Compared to the brief content of the source material stored locally on the smart TV, it expands the amount of plot information and may include all plot details obtained after the search, network-related information, and other information strongly related to the query text.
[0072] Then, the cloud server can input the generated optimized association data, the media data to be played, and the query text into the cloud-based natural language processing and language model application prompt word module to generate prompt words. The prompt words, optimized association data, media data to be played, and query text are then packaged together as parameters for the model and input into the cloud-based companion language model. At this point, the cloud-based companion language model can generate response text data based on the prompt words and the expanded storyline content contained in the optimized association data.
[0073] As one possible implementation, in this embodiment, the cloud server is used to retrieve the full plot data corresponding to the media data to be played in the film and television database; retrieve the network information data corresponding to the query text; and perform vectorized matching on the query text, the full plot data, and the network information data to generate optimized related data with expanded plot information.
[0074] It is easy to understand that the aforementioned full plot data can be the complete plot information corresponding to the media data to be played, stored in a cloud-based film and television database. In this embodiment, the full plot data may include all information data content of the film and television source corresponding to the media data to be played, such as details of the plot not covered by preset plot elements, character backgrounds, behind-the-scenes settings, and production footage. Therefore, the full plot data is more comprehensive than the preset plot elements. Figure 5As shown in this embodiment, after the cloud server receives the media data identifier (such as a movie ID) uploaded by the smart TV, it can call the film and television library data engine to retrieve all information data content corresponding to the film and television source in the cloud film and television library based on the media data identifier. For example, if the media to be played is a historical drama, the full plot data can not only include the main plot corresponding to the film and television source of the media to be played, but also include the detailed background of each character, the historical prototype of the plot scene, the plot adaptation explanation, etc.
[0075] The aforementioned online information data can be publicly available information related to the query text, retrieved by a cloud server through a network information engine. This information could include historical background, industry knowledge, comparisons with similar storylines, and audience interpretations. Figure 5 As shown, the cloud server can use user query text as search keywords to retrieve publicly available online data that is strongly related to the user's conversation. For example, if the query text is "Does the war scene in this historical drama conform to historical facts?", the online information engine can retrieve relevant historical data such as the form of warfare, weaponry, and tactical characteristics of that historical period.
[0076] It is important to understand that in the above vectorized matching process, smart TVs can convert query text, full plot data, and online information data into vector form respectively. Then, by calculating vector similarity, they can filter out information fragments with high similarity to the query text vector to achieve accurate matching.
[0077] For example, the cloud server can use natural language processing technology to convert the query text, the retrieved full-scale plot data, and the online information data into high-dimensional vectors. Then, it calculates the similarity between the full-scale plot data vector and the query text vector, and between the online information data vector and the query text vector, respectively, filtering out information fragments from the full-scale plot data vector and the online information data vector whose similarity to the query text vector is higher than a preset vector threshold (e.g., 0.8). Finally, these filtered information fragments are integrated, and duplicate content and content irrelevant to the media data to be played are removed to generate optimized correlation data with expanded plot information, ensuring that this data is strongly correlated with the query text and the media data to be played. Figure 5 The text corpus shown is related to user conversations.
[0078] In this embodiment, vectorized matching replaces traditional keyword matching, improving the accuracy of information filtering and ensuring that optimized related data is strongly correlated with the query text and plot, reducing redundant information. Furthermore, by combining full plot data with online information data, the relevance of the response to the plot is guaranteed, while also expanding the depth and breadth of information, making the response results more valuable. Therefore, the optimized related data generated in this embodiment can provide high-quality input for the cloud-based companion language model, further improving the accuracy and richness of the response text.
[0079] In addition, such as Figure 5 As shown, the cloud server can also store a corpus model of simulated dialogue data. This corpus model can be constructed to assist in improving the functionality of the accompanying language model on the cloud or device side. It is used to provide data for simulating dialogues in specific vertical domains, thereby optimizing the dialogue generation capabilities of the accompanying language model on the cloud or device side. Simultaneously, the cloud server can store large amounts of question-and-answer corpus data in the following centralized methods to enhance the capabilities of this corpus model: 1) Based on historical dialogue data cleaning and augmentation, high-frequency, typical, and multi-turn dialogue segments are extracted from the dialogue logs, and data augmentation is performed using methods such as synonym replacement, sentence transformation, and insertion of distractors; 2) Generate FAQs (Frequently Asked Questions) based on rule templates. This involves setting some dialogue templates to generate a large number of structured but natural dialogue pairs automatically generated by the Large Language Model (LLM). Given a "seed dialogue" or "intent description", the LLM generates a batch of diverse variant dialogues on the same topic.
[0080] Finally, the cloud server can generate the response results (such as...) Figure 5 The response text output by the cloud-based companion language model is sent to the smart TV via the network. The smart TV can directly use the response result as response text data for subsequent audio conversion.
[0081] Furthermore, since the response text generated in the cloud contains a wealth of information, and smart TVs have limited data processing capabilities, in order to improve data processing efficiency, such as... Figure 5 As shown, the cloud server can also directly convert the response results into audio-text before sending them to the smart TV.
[0082] In this implementation, the massive data storage and retrieval capabilities of the cloud server provide full plot data and network-related information to compensate for the lack of information on the client side; the high-performance cloud-based accompanying language model, combined with optimized related data, can generate information-rich and logically rigorous response results to meet users' complex query needs; the smart TV only needs to upload a small amount of data (movie ID, query text), resulting in low network transmission pressure and no impact on the smoothness of the viewing experience.
[0083] In summary, this embodiment achieves intelligent switching between the terminal-side and cloud-side models by comparing semantic similarity values with preset thresholds. Simple questions (strongly related to the plot and with sufficient information) are answered quickly by the terminal-side companion language model, reducing latency; complex questions (with insufficient information) are supplemented by the cloud-side model to ensure the completeness of the response; the computing pressure of the terminal-side and cloud-side models is reasonably allocated, reducing the burden on the cloud server and improving the overall interaction efficiency.
[0084] This embodiment discloses a method for converting user voice data into query text, performing semantic matching between the query text and preset plot elements to obtain semantic similarity values, determining a target response strategy based on a comparison between a preset similarity threshold and the semantic similarity values, where the target response strategy includes either an end-side companion language model invocation strategy or a cloud-based model invocation strategy, and, if the target response strategy is an end-side companion language model invocation strategy, obtaining at least one information fragment from the preset plot elements whose semantic similarity value is greater than the preset similarity threshold, performing sentiment semantic analysis on the query text to obtain text prompt words, and inputting the text prompt words and information fragments into an end-side companion language model deployed locally to obtain response text data. If the target response strategy is a cloud-based model invocation strategy, uploading the media data to be played and the query text to a cloud server, the cloud server being used to retrieve the full plot data corresponding to the media data to be played from the film and television database; retrieving the network information data corresponding to the query text; performing vectorized matching on the query text, the full plot data, and the network information data to generate optimized association data with expanded plot information, and inputting the optimized association data and the query text into the cloud-based companion language model to generate a response result; and using the response result sent by the cloud server as the response text data.
[0085] This embodiment achieves intelligent switching between the terminal-side and cloud-based models by comparing semantic similarity values with preset thresholds. Simple questions (strongly related to the plot and with sufficient information) are answered quickly by the terminal-side accompanying language model, reducing latency; complex questions (with insufficient information) are supplemented by the cloud-based model to ensure the completeness of the response; the computing pressure of the terminal-side and cloud-based models is reasonably allocated, reducing the burden on the cloud server and improving the overall interaction efficiency.
[0086] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the cross-device collaborative viewing interaction method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.
[0087] This application also provides a cross-device collaborative viewing interaction device, please refer to... Figure 6 , Figure 6 This is a schematic diagram of the module structure of the cross-device collaborative viewing interaction device according to an embodiment of this application. In this embodiment, the cross-device collaborative viewing interaction device is applied to a smart TV, which is connected to an audio output device. The device includes: The data management module T1 is used to obtain the preset plot elements corresponding to the media data to be played selected by the user. The preset plot elements include plot element information divided according to the time dimension. The response processing module T2 is used to match user voice data with preset plot elements and generate response text data based on the matching results. The data distribution module T3 is used to convert response text data into interactive audio data, and distribute the interactive audio data to the audio output device for playback when outputting media data to be played.
[0088] Optionally, in this embodiment, the response processing module T2 is further configured to convert user voice data into query text, and perform semantic matching between the query text and preset plot elements to obtain semantic similarity values; determine a target response strategy based on the numerical comparison results between the preset similarity threshold and the semantic similarity values, wherein the target response strategy includes an end-side companion language model invocation strategy or a cloud-based model invocation strategy; and generate response text data based on the target response strategy and the query text.
[0089] Optionally, in this embodiment, the response processing module T2 is further configured to, when the target response strategy is the end-side companion language model invocation strategy, obtain at least one information fragment with a semantic similarity value greater than a preset similarity threshold from preset plot elements; perform sentiment semantic analysis on the query text to obtain text prompt words; and input the text prompt words and information fragments into the end-side companion language model deployed locally to obtain response text data.
[0090] Optionally, in this embodiment, the response processing module T2 is further configured to upload the media data to be played and the query text to the cloud server when the target response strategy is the cloud model invocation strategy, so that the cloud server generates optimized related data with expanded plot information based on the media data to be played and the query text, and inputs the optimized related data and the query text into the cloud companion language model to generate the response result; and uses the response result sent by the cloud server as the response text data.
[0091] Optionally, in this embodiment, the cloud server is used to retrieve the full plot data corresponding to the media data to be played in the film and television database; retrieve the network information data corresponding to the query text; and perform vectorized matching on the query text, the full plot data, and the network information data to generate optimized related data with expanded plot information.
[0092] Optionally, in this embodiment, the data distribution module T3 is further configured to convert the response text data into audio source text data according to the user-selected role control parameters, and obtain the target role voice data packet from the preset role voice database based on the audio source text data; generate interactive audio data based on the target role voice data packet and the response text data; and distribute the interactive audio data to the audio output device for playback according to the playback timing strategy when outputting the media data to be played, based on the device role identifier corresponding to the audio output device.
[0093] Optionally, in this embodiment, the data distribution module T3 is further configured to broadcast a device discovery request to the audio output device and generate a list of available devices corresponding to the audio output device in response to the device discovery request; the list of available devices includes paired devices and unpaired devices; determine the selected target paired device and / or target unpaired device according to the user's selection instruction for the list of available devices; if there is a target unpaired device, perform pairing identity verification on the target unpaired device and select the target unpaired device that passes the pairing identity verification as the interactive playback terminal; if there is a target paired device, select the target paired device as the interactive playback terminal; and receive user voice data collected by the interactive playback terminal.
[0094] The cross-device collaborative movie-watching interaction device provided in this application, employing the cross-device collaborative movie-watching interaction method in the above embodiments, can solve the technical problems of cross-device collaborative movie-watching interaction. Compared with the prior art, the beneficial effects of the cross-device collaborative movie-watching interaction device provided in this application are the same as the beneficial effects of the cross-device collaborative movie-watching interaction method provided in the above embodiments, and other technical features in the cross-device collaborative movie-watching interaction device are the same as the features disclosed in the methods of the above embodiments, and will not be repeated here.
[0095] This application provides a smart TV, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the cross-device collaborative viewing interaction method in Embodiment 1 above.
[0096] The following is for reference. Figure 7 This document illustrates a structural schematic diagram of a viewing interaction device suitable for implementing cross-device collaboration in the embodiments of this application. Besides smart TVs, the cross-device collaborative viewing interaction device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as desktop computers. Figure 7 The smart TV shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.
[0097] like Figure 7As shown, a smart TV may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to programs stored in read-only memory (ROM) 1002 or programs loaded from storage device 1003 into random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the smart TV. The processing device 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication device 1009 allows the smart TV to communicate wirelessly or wiredly with other devices to exchange data. While the figure shows smart TVs with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.
[0098] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment disclosed in this application includes a cross-device collaborative movie-watching interactive program product, which includes a cross-device collaborative movie-watching interactive program carried on a computer-readable medium, the cross-device collaborative movie-watching interactive program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the cross-device collaborative movie-watching interactive program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the cross-device collaborative movie-watching interactive program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0099] The smart TV provided in this application employs the cross-device collaborative viewing interaction method described in the above embodiments, which can solve the technical problem of cross-device collaborative viewing interaction. Compared with the prior art, the beneficial effects of the smart TV provided in this application are the same as those of the cross-device collaborative viewing interaction method provided in the above embodiments, and other technical features in the cross-device collaborative viewing interaction device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0100] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0101] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0102] This application provides a storage medium having computer-readable program instructions (i.e., a cross-device collaborative viewing interaction program) stored thereon, the computer-readable program instructions being used to execute the cross-device collaborative viewing interaction method in the above embodiments.
[0103] The storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of the storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0104] The aforementioned storage medium may be included in a smart TV; or it may exist independently and not be installed in a smart TV.
[0105] The aforementioned storage medium carries one or more programs, which, when executed by the smart TV, enable cross-device collaborative viewing interaction.
[0106] Cross-device collaborative viewing interaction program code for performing the operations of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and cross-device collaborative interactive viewing programs according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0108] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0109] The readable storage medium provided in this application is a storage medium that stores computer-readable program instructions (i.e., a cross-device collaborative viewing interaction program) for executing the aforementioned cross-device collaborative viewing interaction method, and can solve the technical problem of cross-device collaborative viewing interaction. Compared with the prior art, the beneficial effects of the storage medium provided in this application are the same as the beneficial effects of the cross-device collaborative viewing interaction method provided in the above embodiments, and will not be repeated here.
[0110] The above are only some embodiments of this application and do not limit the scope of the solution of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.
Claims
1. A cross-device collaborative movie-watching interaction method, characterized in that, The method is applied to a smart TV, which is connected to an audio output device; the method includes: Obtain the preset plot elements corresponding to the media data selected by the user to be played; The system receives user voice data collected by the audio output device, performs data matching between the user voice data and the preset plot elements, and generates response text data based on the matching results. The response text data is converted into interactive audio data, and the interactive audio data is distributed to the audio output device for playback when the media data to be played is output.
2. The cross-device collaborative movie-watching interaction method as described in claim 1, characterized in that, The step of matching the user's voice data with the preset plot elements and generating response text data based on the matching results includes: The user's voice data is converted into query text, and the query text is semantically matched with the preset plot elements to obtain a semantic similarity value; The target response strategy is determined based on the numerical comparison result between the preset similarity threshold and the semantic similarity value. The target response strategy includes an end-side companion language model invocation strategy or a cloud-based model invocation strategy. Response text data is generated based on the target response strategy and the query text.
3. The cross-device collaborative movie-watching interaction method as described in claim 2, characterized in that, The step of generating response text data based on the target response strategy and the query text includes: When the target response strategy is the on-device companion language model invocation strategy, at least one information fragment with a semantic similarity value greater than the preset similarity threshold is obtained from the preset plot elements; Sentiment and semantic analysis is performed on the query text to obtain text prompt words; The text prompts and information fragments are input into a locally deployed edge-side companion language model to obtain response text data.
4. The cross-device collaborative movie-watching interaction method as described in claim 2, characterized in that, The smart TV is connected to a cloud server, and the step of generating response text data based on the target response strategy and the query text includes: When the target response strategy is the cloud model invocation strategy, the media data to be played and the query text are uploaded to the cloud server, so that the cloud server generates optimized association data with expanded plot information based on the media data to be played and the query text, and inputs the optimized association data and the query text into the cloud companion language model to generate response results; The response result sent by the cloud server is used as the response text data.
5. The cross-device collaborative movie-watching interaction method as described in claim 4, characterized in that, The cloud server is used to retrieve the full plot data corresponding to the media data to be played in the film and television database; retrieve the network information data corresponding to the query text; and perform vectorized matching on the query text, the full plot data, and the network information data to generate optimized association data with expanded plot information.
6. The cross-device collaborative movie-watching interaction method as described in claim 1, characterized in that, The step of converting the response text data into interactive audio data and distributing the interactive audio data to the audio output device for playback when outputting the media data to be played includes: The response text data is converted into audio source text data according to the user-selected role control parameters, and the target role voice data package is obtained from the preset role voice database based on the audio source text data. Generate interactive audio data based on the target character's voice data packet and the response text data; Based on the device role identifier corresponding to the audio output device, when outputting the media data to be played, the interactive audio data is distributed to the audio output device for playback according to the playback timing strategy.
7. The cross-device collaborative movie-watching interaction method as described in claim 1, characterized in that, The step of receiving user voice data collected by the audio output device includes: A device discovery request is broadcast to the audio output device, and a list of available devices corresponding to the audio output device in response to the device discovery request is generated; the list of available devices includes paired devices and unpaired devices; The selected target paired device and / or target unpaired device are determined based on the user's selection instruction for the list of available devices; In the case of the existence of the target unpaired device, the pairing identity verification is performed on the target unpaired device, and the target unpaired device that passes the pairing identity verification is selected as the interactive playback terminal; If the target paired device exists, the target paired device will be selected as the interactive playback terminal; Receive user voice data collected by the interactive playback terminal.
8. A cross-device collaborative movie-watching interactive device, characterized in that, The device is used in a smart TV, the smart TV being connected to an audio output device, and the device includes: The data management module is used to obtain the preset plot elements corresponding to the media data selected by the user to be played; The response processing module is used to receive user voice data collected by the audio output device, perform data matching between the user voice data and the preset plot elements, and generate response text data based on the matching results. The data distribution module is used to convert the response text data into interactive audio data, and distribute the interactive audio data to the audio output device for playback when outputting the media data to be played.
9. A smart TV, characterized in that, The smart TV includes: a memory, a processor, and a cross-device collaborative viewing interaction program stored on the memory and executable on the processor, the cross-device collaborative viewing interaction program being configured to implement the steps of the cross-device collaborative viewing interaction method as described in any one of claims 1 to 7.
10. A storage medium, characterized in that, The storage medium stores a cross-device collaborative movie-watching interaction program, which, when executed by a processor, implements the steps of the cross-device collaborative movie-watching interaction method as described in any one of claims 1 to 7.