Video interaction method, and apparatus

By introducing a plot interaction mode during video playback, users can pause playback at key frames and get relevant questions and answers, solving the problem of understanding the plot during the viewing process and improving the user's viewing experience.

WO2026021142A1PCT designated stage Publication Date: 2026-01-29HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/104351
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-25
Filing Date
2025-06-27
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

In existing technologies, users find it difficult to efficiently understand the plot of film and television works during the viewing process, and there is a lack of effective interactive methods to improve the viewing experience.

Method used

During video playback, the interactive story mode pauses playback and displays questions and answers related to the storyline. Users can select questions and obtain answers to learn about story details and the creators' intentions.

Benefits of technology

Through interactive Q&A sessions related to the plot, users can stay informed about the plot development and the creator's intentions, enhancing their viewing experience and avoiding confusion caused by too many plot branches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025104351_29012026_PF_FP_ABST
    Figure CN2025104351_29012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a video interaction method, and an apparatus. The video interaction method of the present application comprises: upon determining entry into a plot interaction mode, and when a playback reaches a plot interaction frame of a first video, pausing the playback of the first video, and displaying a first question on a playback interface, the first question being related to a plot expressed by the plot interaction frame (401); and upon receiving a first instruction corresponding to the first question triggered by a user, in response to the first instruction, displaying a first answer on the playback interface, the first answer being used for answering the first question, so as to help the user understand the plot expressed by the plot interaction frame (402). The present application helps users to better understand plots, thereby improving the viewing experience of users.
Need to check novelty before this filing date? Find Prior Art

Description

Video interaction methods and devices Technical Field

[0001] This application relates to the field of video playback, and more particularly to a video interaction method and apparatus. Background Technology

[0002] Film and television works, as an art form, use images and sound to express emotions and convey information. Directors edit the footage to create a film, which is then presented to the user in a sequential frame-by-frame manner, incorporating flashbacks, intercuts, and other techniques. While watching, users can understand the plot by combining their own interpretation with their viewing experience. Furthermore, they can engage in interactive discussions on social media, read professional film reviews, and other external supplementary information to help them grasp the director's intentions and the details of the actors' performances. This allows them to later revisit classic scenes and rediscover plot points and highlights missed during the initial viewing, thus enhancing the overall viewing experience. Therefore, finding more efficient interactive methods to help users understand the plot is a crucial aspect of improving the viewing experience. Summary of the Invention

[0003] This application provides a video interaction method and device to help users better understand the plot and enhance their viewing experience.

[0004] In a first aspect, this application provides a video interaction method, comprising: in a story interaction mode, when playing a story interaction frame of a first video, pausing the playback of the first video and displaying a first question on the playback interface, the first question being related to the story expressed by the story interaction frame; and, in response to receiving a first instruction triggered by a user corresponding to the first question, displaying a first answer on the playback interface, the first answer being used to answer the first question to help the user understand the story expressed by the story interaction frame.

[0005] This application generates plot-related Q&A (including questions and answers) in the video and the plot interaction frames. When the plot interaction frames are played, the Q&A format allows the user to interact with the plot, making the audience more immersive in the plot and the characters' situations. This not only allows them to understand the plot in a timely manner, but also allows them to capture the information that the creators of the film and television works want to convey, thereby better understanding the plot and improving the user's viewing experience.

[0006] The narrative interaction mode is a video interaction method provided in this application during video playback. After determining to enter narrative interaction mode, when the first video (e.g., a film, short video, etc.) reaches a narrative interaction frame, the user can interact with the plot-related information through a narrative Q&A format (including questions and answers). The user can select a question they want to know (the first question), which can be related to the plot expressed in the narrative interaction frame, such as the character's personality or motivation in the frame, and then obtain a corresponding answer (the first answer). This allows the user to understand the plot promptly, capture the information the creators of the first video want to convey, and thus better comprehend the story.

[0007] The method for determining whether to enter the story interaction mode in this application may include: before determining whether to enter the story interaction mode, when the playback reaches the scheduled interaction frame of the first video, displaying a prompt message to prompt the user to confirm whether to enter the story interaction mode, wherein the scheduled interaction frame is n seconds earlier than the story interaction frame, where n≥1. After this, two branches can occur:

[0008] 1) Upon receiving a fourth command triggered by the user, the system responds to the fourth command by determining whether to enter the story interaction mode. This fourth command includes instructions generated by the user operating the remote control or clicking a control on the touchscreen to indicate confirmation.

[0009] 2) Upon receiving the fifth instruction triggered by the user, in response to the fifth instruction, determine not to enter the story interaction mode and continue playing the first video. This fifth instruction includes instructions generated by the user operating the remote control or clicking a control on the touchscreen to indicate negation.

[0010] A scheduled interaction frame is pre-set in the first video. This scheduled interaction frame is earlier than the story interaction frame. For example, the scheduled interaction frame is 5 seconds earlier than the story interaction frame. In a 24-frame video, the scheduled interaction frame is about 120 frames earlier than the story interaction frame.

[0011] When the scheduled interactive frame is played, a prompt message can be displayed on the playback interface to ask the user to confirm whether to enter the interactive story mode. For example, a bubble-like prompt might appear on the playback control bar, displaying the message: "High-energy content ahead! Press the up button to enter interactive story mode, press the down button to not enter interactive story mode." The up or down button interaction can be tailored to a TV or set-top box scenario. Clicking the up button on the TV or set-top box remote triggers the process of entering interactive story mode, while clicking the down button prevents entry. Therefore, the fourth instruction can be generated by the user clicking the up button on the remote, and the fifth instruction can be generated by the user clicking the down button. Other interaction methods can also be used, such as confirmation or rejection controls on a touchscreen, allowing the user to select whether to enter or not enter interactive story mode. This application does not specifically limit the interaction method.

[0012] Optionally, if no user-triggered instruction is received after a preset waiting period, the system will not enter the story interaction mode and will continue playing the first video.

[0013] If the user neither clicks the confirmation control nor the rejection control when the appointment interaction frame is played, it can be determined that the story interaction mode will not be entered after waiting for a certain period of time (this period is preset, for example, 2 seconds).

[0014] After the process of determining whether to enter the story interaction mode, if it is determined not to enter the story interaction mode, the first video will continue to play, and there will be no interaction with the user at the story interaction frames.

[0015] Interactive story frames serve as user interfaces. Questions and answers (including questions and answers) related to the storyline expressed in these frames can be obtained through offline preprocessing on the video platform, then transmitted to the electronic device via a communication link and saved. The process of the video platform generating questions and answers related to the storyline expressed in the interactive story frames offline may include:

[0016] 1) Obtain candidate frames of the first video, which are used to represent shots of the first video;

[0017] 2) Perform facial recognition on each candidate frame to obtain interactive plot frames, which include characters in the play facing the audience and occupying a central position.

[0018] 3) Obtain plot-related information from the first video and generate a plot description text;

[0019] 4) Align the interactive story frames with the story description text to obtain at least one set of questions and answers (including questions and answers) for the interactive story frames.

[0020] Based on the above process, at least one set of question-and-answer text can be obtained from the plot introduction text. This text is saved so that when the plot interaction frame is played, the corresponding set of question-and-answer text can be extracted and displayed.

[0021] Furthermore, building upon the offline preprocessing steps of the aforementioned video platforms, it's possible to add a feature where text-based dialogue questions are used to generate videos based on the characters' images and voiceprints. These videos can then be used as video material in the interactive dialogue mode and persistently saved on the video platform. This process generates at least one set of questions and answers based on the plot summary text.

[0022] In the interactive story mode, when the playback reaches a story interaction frame of the first video, the first video is paused, and the first question is displayed on the playback interface. At this time, the background of the playback interface can be the story interaction frame. Based on the aforementioned offline generation process of questions and answers related to the story expressed by the story interaction frame, the first question can be displayed in the following two ways:

[0023] (1) Display the first question in text form on the playback interface.

[0024] The text for the first question comes from at least one set of question-and-answer text obtained from the offline preprocessing process of the aforementioned video platform. One question per interactive plot frame can be displayed sequentially on the playback interface, or multiple questions corresponding to interactive plot frames can be displayed simultaneously; there is no specific limitation on this. Optionally, the interactive plot frame can be used as the background at this time.

[0025] When the playback interface displays the text of a question, the first question mentioned above is that question; or, when the playback interface displays the text of multiple questions, the first question can be one of the multiple questions. The user can use the remote control or touch screen to tap the text of the first question to generate a first instruction, which indicates the question selected by the user (i.e., the first question) or indicates that the user chooses to answer the currently displayed question (the first question).

[0026] (2) Play the video of the first question on the playback interface. The video of the first question includes the target character uttering the first question. The target character is a character in the play who faces the audience and occupies a central position in the interactive plot frame. The video of the first question comes from at least one set of question-and-answer videos obtained from the offline preprocessing process of the aforementioned video platform. The video of one question (the first question) in the interactive plot frame can be played one at a time on the playback interface in a turn-by-turn manner. The user can operate the remote control or touch screen to click the control indicating confirmation in the video of the first question to generate a first instruction. The first instruction instructs the user to select to answer the currently played question (the first question). Optionally, the aforementioned target character can be presented as a digital human. The digital human image of the target character can be generated using 3D image technology, without specific limitations.

[0027] The first answer may include at least one of the plot summary or film review of the first video, or the first answer may be based on at least one of the plot summary or film review of the first video, so that the user can understand the plot of the first video through the first answer, including the plot expressed by the plot interaction frame and the frame before it.

[0028] Corresponding to the first question above, the first answer can also be displayed in the following two ways:

[0029] (1) Display the first answer in text form on the playback interface. The text of the first answer comes from at least one set of question and answer text obtained from the offline preprocessing process of the aforementioned video platform. When the user selects the first question or chooses to answer the currently displayed first question, the text of the corresponding answer (the first answer) can be extracted and displayed on the playback interface.

[0030] (2) Play the video of the first answer on the playback interface. The video of the first answer includes the target character's oral first answer. The target character is a character in the play who faces the audience and occupies a central position in the interactive plot frame. The video of the first answer comes from at least one set of question-and-answer videos obtained from the offline preprocessing process of the aforementioned video platform. When the user selects to answer the first question currently being played, the corresponding answer (the first answer) video can be extracted and played on the playback interface.

[0031] In one possible implementation, after reading the first answer—whether the text or the video—the user can choose to view the next question (the second question) and its answer (the second answer). That is,

[0032] The system receives a second instruction triggered by the user, corresponding to a second question. This second question is related to the plot expressed in the interactive plot frame and is different from the first question mentioned above. In response to the second instruction, the system displays the second question on the playback interface. The aforementioned second instruction includes an instruction generated by the user operating a remote control or touchscreen to click on the text of the second question; or, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to click on a control indicating the next question.

[0033] In one possible implementation, after viewing the Nth answer (either the text or the video), the user can choose to stop watching the Q&A segments of the interactive storyline and continue watching the first video, where N ≥ 1.

[0034] The system receives a third command triggered by the user; in response to the third command, it exits the story interaction mode and resumes playback of the first video from the story interaction frame. The aforementioned third command includes commands generated by the user operating the remote control or clicking the control on the touchscreen to indicate returning. After the user chooses not to continue watching the Q&A in the story interaction frame, playback of the first video can resume from the frame (story interaction frame) where the original video (first video) was interrupted. At this point, it does not exit the story interaction mode of the first video, but rather stops displaying the undisplayed Q&A in the story interaction frame. That is, after watching the Nth answer, even if there are still unwatched Q&A in that story interaction frame, the user can still trigger the third command, and based on this third command, the undisplayed Q&A in the story interaction frame can be stopped, and playback of the first video can continue.

[0035] Optionally, if all questions and answers in the story interaction frame have been displayed after the Nth answer has been shown, the story interaction mode of that story interaction frame can be automatically exited and the first video can continue to play. In this case, no user-triggered command is required.

[0036] In one possible implementation, after viewing the Nth answer (either the text or the video), the user can choose to exit the interactive story mode and continue watching the first video. Even if a new interactive story frame is encountered during subsequent playback, the question and answer for that frame will not be displayed again, where N≥1. That is,

[0037] The system receives a sixth command triggered by the user; in response to the sixth command, it exits the story interaction mode and resumes playing the first video. The aforementioned sixth command includes commands generated by the user operating the remote control or tapping the exit control on the touchscreen. After the user selects to exit the story interaction mode, playback of the first video can resume from the frame (story interaction frame) where the original video (first video) was interrupted. At this point, the story interaction mode for the first video has exited; even if there is another story interaction frame, the questions and answers for the next story interaction frame will not be displayed when playback reaches that frame; instead, the first video will continue playing.

[0038] It is clear that the interactive story mode does not change the first video itself. Users can still watch the first video in its entirety without experiencing the confusion that can occur due to too many story branches, as is the case with related technologies.

[0039] In this embodiment, a story-driven interactive mode is implemented. After confirming the entry into this mode, when the first video plays a story-driven interactive frame, the user can select the questions they want to know and receive corresponding answers. Through such story-related Q&A, users can quickly understand the plot's direction, the character settings and motives of the target characters, and detailed analysis, without needing to search websites or forums. This allows users to understand the information the creators want to convey, thus better comprehending the plot and enhancing the user's viewing experience.

[0040] Secondly, this application provides a video interaction device, comprising: a display module, configured to, after determining that a story interaction mode is to be entered, pause playback of the first video when the playback reaches a story interaction frame of the first video, and display a first question on the playback interface, the first question being related to the storyline expressed by the story interaction frame; and, in response to receiving a first instruction triggered by a user corresponding to the first question, display a first answer on the playback interface, the first answer being used to answer the first question to help the user understand the storyline expressed by the story interaction frame.

[0041] In one possible implementation, the system further includes: a processing module for receiving a second instruction triggered by a user corresponding to a second question, the second question being related to the plot expressed by the plot interaction frame and different from the first question; and a display module for responding to the second instruction by displaying the second question on the playback interface.

[0042] In one possible implementation, the processing module is further configured to receive a third instruction triggered by the user; the display module is further configured to continue playing the first video in response to the third instruction.

[0043] In one possible implementation, the display module is further configured to display a prompt message when the first video is played to a scheduled interaction frame before entering the story interaction mode. The prompt message is used to prompt the user to confirm whether to enter the story interaction mode. The scheduled interaction frame is n seconds earlier than the story interaction frame, where n≥1. The processing module is further configured to determine whether to enter the story interaction mode in response to the fourth instruction triggered by the user.

[0044] In one possible implementation, the processing module is further configured to, upon receiving a fifth instruction triggered by the user, determine, in response to the fifth instruction, not to enter the story interaction mode and continue playing the first video.

[0045] In one possible implementation, the processing module is further configured to, after waiting for a preset time, if no instruction triggered by the user is received, determine not to enter the story interaction mode and continue playing the first video.

[0046] In one possible implementation, the display module is specifically used to display the first question in text form on the playback interface.

[0047] In one possible implementation, the first instruction includes instructions generated by the user operating a remote control or touchscreen to tap the text of the first question.

[0048] In one possible implementation, the display module is specifically used to play a video of the first question on the playback interface. The video of the first question includes a target person uttering the first question. The target person is a character in the play who is facing the audience and occupies a central position in the interactive plot frame.

[0049] In one possible implementation, the first instruction includes an instruction generated by the user operating a remote control or touchscreen to click a control in the video of the first question that indicates confirmation.

[0050] In one possible implementation, the display module is specifically used to display the first answer in text form on the playback interface; or, to play a video of the first answer on the playback interface, wherein the video of the first answer includes a target person dictating the first answer, and the target person is a character in the play who is facing the audience and occupies a central position in the interactive plot frame.

[0051] In one possible implementation, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to tap the text of the second question; or, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to tap a control indicating the next question.

[0052] In one possible implementation, the third instruction includes instructions generated by the user operating a remote control or a touchscreen to tap a control indicating a return.

[0053] In one possible implementation, the fourth instruction includes instructions generated by the user operating a remote control or touchscreen to click on a control indicating confirmation.

[0054] In one possible implementation, the fifth instruction includes an instruction generated by the user operating a remote control or touchscreen to click a control that indicates negation.

[0055] In one possible implementation, the first answer includes at least one of the plot summary or film review of the first video; or, the first answer is obtained based on at least one of the plot summary or film review of the first video.

[0056] Thirdly, this application provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to perform the method as described in any one of the first aspects above.

[0057] Fourthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the first aspects above.

[0058] Fifthly, this application provides a computer program product comprising computer program code, which, when run on a computer, causes the computer to perform the method described in any one of the first aspects above. Attached Figure Description

[0059] Figure 1a is a schematic diagram of a movie-watching scenario in this application;

[0060] Figure 1b is an exemplary schematic diagram of the movie viewing system of this application;

[0061] Figure 2 is an exemplary structural diagram of the electronic device 200 of this application;

[0062] Figure 3 is an exemplary software structure block diagram of the electronic device 200 provided in an embodiment of this application;

[0063] Figure 4 is a flowchart of process 400 of the video interaction method provided in this application;

[0064] Figure 5 is a schematic diagram of the playback interface of this application;

[0065] Figure 6 is a flowchart illustrating the video interaction method of this application;

[0066] Figure 7 is a flowchart illustrating the video interaction method of this application;

[0067] Figure 8 is a structural schematic diagram of the video interaction device 800 provided in this application;

[0068] Figure 9 shows a schematic block diagram of an apparatus 900 according to an embodiment of this application. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0070] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0071] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0072] Figure 1a is a schematic diagram of the viewing scenario of this application. As shown in Figure 1a, this is a viewing scenario in a home environment, where the user plays the film or television works (including movies, TV shows, short videos, etc.) they want to watch through a television set. The film or television works can be provided by a video platform (e.g., Huawei Video), and the corresponding client is a video application (APP) installed on the television (e.g., Huawei Video APP), which can provide the user with an interactive interface and a playback window.

[0073] When a user interacts with the interface provided by the video app, it triggers a connection to a video or movie stored in the storage space provided by the video platform. The video or movie is then cached on the network to the local storage space corresponding to the video app and then played on the TV screen.

[0074] In this application, user operations can be input via a TV remote control, a TV touchscreen, voice commands, or a smart terminal connected to the TV (such as a smartphone), without specific limitations. The caching method for video content aims to ensure clear and smooth playback, and its specific implementation is not limited.

[0075] It should be noted that, in addition to the home viewing scenario described in Figure 1a, this application may also include viewing scenarios of watching film and television works in the home or other environments via mobile phones, tablets, or other electronic devices, without making specific limitations.

[0076] Figure 1b is an exemplary schematic diagram of the movie-watching system of this application. As shown in Figure 1b, the movie-watching system includes a client and a video platform, wherein...

[0077] The client can be a video app installed on an electronic device, as described above, used to provide interactive interfaces, local caching, and related video processing. The video platform can be the server or cloud of the video provider corresponding to the aforementioned video app, used to provide film and television works and related video processing.

[0078] The network between the client and the video platform can be a communication network that supports short-range communication technologies, such as a communication network that supports Wireless-Fidelity (WIFI) technology; or a communication network that supports Bluetooth technology; or a communication network that supports Near Field Communication (NFC) technology; etc.; or the communication network can also be a communication network that supports fourth-generation (4G) access technology, such as Long Term Evolution (LTE) access technology; or the communication network can also be a communication network that supports fifth-generation (5G) access technology, such as New Radio (NR) access technology; or the communication network can also be a communication network that supports third-generation (3G) access technology, such as Universal Mobile Telecommunications System (UMTS) access technology; or the communication network can also be a communication network that supports multiple wireless technologies, such as a communication network that supports LTE and NR technologies; or the communication network can also be adapted to future-oriented communication technologies, which are not specifically limited in this application.

[0079] Figure 2 is an exemplary structural diagram of the electronic device 200 of this application. As shown in Figure 2, the electronic device 200 may include any one of the following electronic devices: mobile phone, tablet computer, laptop computer, smart screen, wearable device, in-vehicle device, augmented reality (AR) / virtual reality (VR) device, etc.

[0080] It should be understood that the electronic device 200 shown in Figure 2 is merely an example, and the electronic device 200 may have more or fewer components than shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in Figure 2 may be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application-specific integrated circuits.

[0081] Electronic device 200 may include: processor 210, external memory interface 220, internal memory 221, universal serial bus (USB) interface 230, charging management module 240, power management module 241, battery 242, antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, sensor module 280, button 290, motor 291, indicator 292, camera 293, display screen 294, and subscriber identification module (SIM) card interface 295, etc. The sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an accelerometer sensor 280E, a distance sensor 280F, a proximity sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0082] Processor 210 may include one or more processing units, such as: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, memory, video codec, digital signal processor (DSP), baseband processor, and / or neural network processing unit (NPU), etc. Different processing units may be independent devices or integrated into one or more processors.

[0083] The controller can be the nerve center and command center of the electronic device 200. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0084] The processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can store instructions or data that the processor 210 has just used or that are used repeatedly. If the processor 210 needs to use the instruction or data again, it can directly retrieve it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.

[0085] In some embodiments, the processor 210 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0086] The I2C interface is a bidirectional synchronous serial bus, including a serial data line (SDA) and a serial clock line (SCL). In some embodiments, the processor 210 may include multiple I2C buses. The processor 210 can couple to the touch sensor 280K, charger, flash, camera 293, etc., through different I2C bus interfaces. For example, the processor 210 can couple to the touch sensor 280K through the I2C interface, enabling the processor 210 and the touch sensor 280K to communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 200.

[0087] The I2S interface can be used for audio communication. In some embodiments, the processor 210 may include multiple I2S buses. The processor 210 can be coupled to the audio module 270 via the I2S bus to enable communication between the processor 210 and the audio module 270. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the I2S interface to enable the function of answering phone calls through a Bluetooth headset.

[0088] The PCM interface can also be used for audio communication, sampling, quantizing, and encoding analog signals. In some embodiments, the audio module 270 and the wireless communication module 260 can be coupled via the PCM bus interface. In some embodiments, the audio module 270 can also transmit audio signals to the wireless communication module 260 via the PCM interface, enabling the function of answering phone calls through a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0089] The UART interface is a universal serial data bus used for asynchronous communication. This bus can be a bidirectional communication bus. It converts the data to be transmitted between serial and parallel communication. In some embodiments, the UART interface is typically used to connect the processor 210 and the wireless communication module 260. For example, the processor 210 communicates with the Bluetooth module in the wireless communication module 260 via the UART interface to implement Bluetooth functionality. In some embodiments, the audio module 270 can transmit audio signals to the wireless communication module 260 via the UART interface to enable music playback through Bluetooth headphones.

[0090] The MIPI interface can be used to connect the processor 210 to peripheral devices such as the display screen 294 and the camera 293. The MIPI interface includes a camera serial interface (CSI) and a display serial interface (DSI). In some embodiments, the processor 210 and the camera 293 communicate via the CSI interface to enable the electronic device 200 to capture images. The processor 210 and the display screen 294 communicate via the DSI interface to enable the electronic device 200 to display images.

[0091] The GPIO interface can be configured via software. It can be configured as a control signal or a data signal. In some embodiments, the GPIO interface can be used to connect the processor 210 to a camera 293, a display screen 294, a wireless communication module 260, an audio module 270, a sensor module 280, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0092] USB port 230 is a USB standard compliant interface, specifically a Mini USB port, Micro USB port, or USB Type-C port. USB port 230 can be used to connect a charger to charge electronic device 200, and can also be used for data transfer between electronic device 200 and peripheral devices. It can also be used to connect headphones for audio playback. This interface can also be used to connect other electronic devices, such as AR devices.

[0093] It should be understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0094] The charging management module 240 receives charging input from a charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 receives charging input from the wired charger via a USB interface 230. In some wireless charging embodiments, the charging management module 240 receives wireless charging input via the wireless charging coil of the electronic device 200. While charging the battery 242, the charging management module 240 can also supply power to the electronic device via the power management module 241.

[0095] The power management module 241 connects the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240, providing power to the processor 210, internal memory 221, external memory, display screen 294, camera 293, and wireless communication module 260. The power management module 241 can also monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage current, impedance). In some other embodiments, the power management module 241 may also be located within the processor 210. In other embodiments, the power management module 241 and the charging management module 240 may be housed in the same device.

[0096] The wireless communication function of electronic device 200 can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor, and baseband processor.

[0097] Antenna 1 and antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 200 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0098] The mobile communication module 250 can provide solutions for wireless communication, including 2G / 3G / 4G / 5G, applied to the electronic device 200. The mobile communication module 250 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 1. In some embodiments, at least some functional modules of the mobile communication module 250 may be housed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 may be housed in the same device.

[0099] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 270A, receiver 270B, etc.) or displays images or videos through the display screen 294. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 210 and may be housed in the same device as the mobile communication module 250 or other functional modules.

[0100] The wireless communication module 260 can provide solutions for wireless communication applications on the electronic device 200, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 260 can be one or more devices integrating at least one communication processing module. The wireless communication module 260 receives electromagnetic waves via antenna 2, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 210. The wireless communication module 260 can also receive signals to be transmitted from processor 210, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0101] In some embodiments, antenna 1 of electronic device 200 is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, enabling electronic device 200 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0102] Electronic device 200 implements display functions through a GPU, a display screen 294, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 210 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0103] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, electronic device 200 may include one or N displays 294, where N is a positive integer greater than 1.

[0104] Electronic device 200 can perform shooting functions through ISP, camera 293, video codec, GPU, display screen 294 and application processor.

[0105] The ISP (Image Signal Processor) is used to process data fed back from the camera 293. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 293.

[0106] Camera 293 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the electronic device 200 may include one or N cameras 293, where N is a positive integer greater than 1.

[0107] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when electronic device 200 selects a frequency, the DSP is used to perform Fourier transforms on the frequency energy.

[0108] Video codecs are used to compress or decompress digital video. Electronic device 200 may support one or more video codecs. Thus, electronic device 200 can play or record video in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.

[0109] An NPU (Neural Processing Unit) is a neural network (NN) computing processor that, by borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, rapidly processes input information and can continuously learn on its own. NPUs enable intelligent cognitive applications in electronic devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0110] The external storage interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 200. The external memory card communicates with the processor 210 through the external storage interface 220 to perform data storage functions. For example, music, video, and other files can be saved on the external memory card.

[0111] Internal memory 221 can be used to store computer executable program code, which includes instructions. Processor 210 executes various functional applications and data processing of electronic device 200 by running the instructions stored in internal memory 221. Internal memory 221 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback, image playback, etc.), etc. The data storage area may store data created during the use of electronic device 200 (such as audio data, phonebook, etc.). Furthermore, internal memory 221 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0112] Electronic device 200 can implement audio functions such as music playback and recording through audio module 270, speaker 270A, receiver 270B, microphone 270C, headphone jack 270D, and application processor.

[0113] The audio module 270 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 may be located in the processor 210, or some functional modules of the audio module 270 may be located in the processor 210.

[0114] The speaker 270A, also known as a "loudspeaker," is used to convert audio electrical signals into sound signals. The electronic device 200 can listen to music or make hands-free calls through the speaker 270A.

[0115] The receiver 270B, also known as the "earpiece," is used to convert audio electrical signals into sound signals. When the electronic device 200 answers a telephone call or voice message, the receiver 270B can be brought close to the ear to listen to the voice.

[0116] Microphone 270C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 270C, inputting the sound signal into microphone 270C. Electronic device 200 may have at least one microphone 270C. In some embodiments, electronic device 200 may have two microphones 270C, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, electronic device 200 may also have three, four, or more microphones 270C, which can collect sound signals, reduce noise, identify the sound source, and perform directional recording, etc.

[0117] The headphone jack 270D is used to connect wired headphones. The headphone jack 270D can be a USB 230 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0118] Pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 280A can be disposed on display screen 294. There are many types of pressure sensors 280A, such as resistive pressure sensors, inductive pressure sensors, and capacitive pressure sensors. A capacitive pressure sensor may include at least two parallel plates with conductive material. When force is applied to pressure sensor 280A, the capacitance between the electrodes changes. Electronic device 200 determines the pressure intensity based on the change in capacitance. When a touch operation is applied to display screen 294, electronic device 200 detects the intensity of the touch operation based on pressure sensor 280A. Electronic device 200 can also calculate the touch position based on the detection signal from pressure sensor 280A. In some embodiments, touch operations applied to the same touch position but with different touch operation intensities can correspond to different operation commands. For example, when a touch operation with an intensity less than a first pressure threshold is applied to the SMS application icon, a command to view an SMS is executed. When a touch operation with an intensity greater than or equal to the first pressure threshold is applied to the SMS application icon, a command to create a new SMS is executed.

[0119] The gyroscope sensor 280B can be used to determine the motion attitude of the electronic device 200. In some embodiments, the gyroscope sensor 280B can determine the angular velocity of the electronic device 200 around three axes (i.e., the x, y, and z axes). The gyroscope sensor 280B can be used for image stabilization. For example, when the shutter is pressed, the gyroscope sensor 280B detects the angle of the electronic device 200's shake, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to counteract the shake of the electronic device 200 through reverse movement, thus achieving image stabilization. The gyroscope sensor 280B can also be used in navigation and motion-sensing game scenarios.

[0120] The barometric pressure sensor 280C is used to measure air pressure. In some embodiments, the electronic device 200 calculates altitude using the air pressure value measured by the barometric pressure sensor 280C to assist in positioning and navigation.

[0121] The magnetic sensor 280D includes a Hall sensor. The electronic device 200 can use the magnetic sensor 280D to detect the opening and closing of the flip cover. In some embodiments, when the electronic device 200 is a flip phone, the electronic device 200 can detect the opening and closing of the flip cover using the magnetic sensor 280D. Then, based on the detected opening and closing state of the cover or the flip cover, features such as automatic flip unlocking can be set.

[0122] The accelerometer 280E can detect the magnitude of acceleration of electronic device 200 in various directions (typically three axes). When electronic device 200 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of electronic device, and can be applied to applications such as screen orientation switching and pedometers.

[0123] A distance sensor 280F is used to measure distance. Electronic device 200 can measure distance via infrared or laser. In some embodiments, during a shooting scene, electronic device 200 can utilize the distance sensor 280F to measure distance for rapid focusing.

[0124] The proximity sensor 280G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The LED may be an infrared LED. The electronic device 200 emits infrared light outward through the LED. The electronic device 200 uses the photodiode to detect infrared reflected light from surrounding objects. When sufficient reflected light is detected, it can be determined that there is an object around the electronic device 200. When insufficient reflected light is detected, the electronic device 200 can determine that there is no object around it. The electronic device 200 may use the proximity sensor 280G to detect when a user holds the electronic device 200 close to their ear for a call, so as to automatically turn off the screen to save power. The proximity sensor 280G can also be used in holster mode and pocket mode for automatic unlocking and locking of the screen.

[0125] The ambient light sensor 280L is used to sense ambient light intensity. The electronic device 200 can adaptively adjust the brightness of its display screen 294 based on the sensed ambient light intensity. The ambient light sensor 280L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 280L can also work in conjunction with the proximity sensor 280G to detect whether the electronic device 200 is in a pocket, preventing accidental touches.

[0126] The fingerprint sensor 280H is used to collect fingerprints. The electronic device 200 can utilize the characteristics of the collected fingerprints to achieve fingerprint unlocking, accessing application locks, taking photos with fingerprints, answering calls with fingerprints, etc.

[0127] Temperature sensor 280J is used to detect temperature. In some embodiments, electronic device 200 uses the temperature detected by temperature sensor 280J to execute a temperature handling strategy. For example, when the temperature reported by temperature sensor 280J exceeds a threshold, electronic device 200 performs thermal protection by reducing the performance of the processor located around temperature sensor 280J to reduce power consumption. In other embodiments, when the temperature is below another threshold, electronic device 200 heats battery 242 to prevent abnormal shutdown of electronic device 200 due to low temperature. In still other embodiments, when the temperature is below yet another threshold, electronic device 200 boosts the output voltage of battery 242 to prevent abnormal shutdown due to low temperature.

[0128] Touch sensor 280K, also known as a "touch panel," can be located on display screen 294. The touch sensor 280K and display screen 294 together form a touchscreen, also known as a "touchscreen." Touch sensor 280K detects touch operations applied to or around it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 294. In other embodiments, touch sensor 280K may also be located on the surface of electronic device 200, in a different position than display screen 294.

[0129] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals from the vibrating bone segments of the human vocal cords. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 280M can also be incorporated into headphones to form bone conduction headphones. The audio module 270 can parse the voice signals from the vibrating bone segments of the vocal cords acquired by the bone conduction sensor 280M to realize voice functionality. The application processor can parse heart rate information from the blood pressure signals acquired by the bone conduction sensor 280M to realize heart rate detection functionality.

[0130] Buttons 290 include a power button, volume buttons, etc. Buttons 290 can be mechanical buttons or touch-sensitive buttons. Electronic device 200 can receive button input and generate key signal inputs related to user settings and function control of electronic device 200.

[0131] Motor 291 can generate vibration alerts. Motor 291 can be used for incoming call vibration alerts or for touch vibration feedback. For example, different vibration feedback effects can be corresponding to touch operations applied to different applications (such as taking photos, playing audio, etc.). Motor 291 can also correspond to different vibration feedback effects for touch operations applied to different areas of the display screen 294. Different application scenarios (such as time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also be customized.

[0132] Indicator 292 can be an indicator light, which can be used to indicate charging status, power changes, messages, missed calls, notifications, etc.

[0133] The SIM card interface 295 is used to connect a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to make contact with and separate from the electronic device 200. The electronic device 200 can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 295 simultaneously. The multiple cards can be of the same or different types. The SIM card interface 295 is also compatible with different types of SIM cards. The SIM card interface 295 is also compatible with external memory cards. The electronic device 200 interacts with the network through the SIM card to realize functions such as calls and data communication. In some embodiments, the electronic device 200 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 200 and cannot be separated from the electronic device 200.

[0134] The software system of electronic device 200 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses the layered architecture Android system as an example to exemplify the software structure of electronic device 200.

[0135] Figure 3 is an exemplary software structure block diagram of an electronic device 200 provided in an embodiment of this application.

[0136] The layered architecture of the electronic device 200 divides the software into several layers, each with a clear role and division of labor. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.

[0137] The application layer can include a series of application packages.

[0138] As shown in Figure 3, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, and SMS.

[0139] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0140] As shown in Figure 3, the application framework layer may include a window manager, a phone manager, a content provider, a view system, a resource manager, a notification manager, etc.

[0141] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0142] The phone manager is used to provide communication functions for electronic devices 200. For example, it manages call status (including connection and disconnection).

[0143] Content providers store and retrieve data, making that data accessible to applications. This data may include videos, images, audio, made and received phone calls, browsing history and bookmarks, phone books, etc.

[0144] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0145] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0146] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0147] The Android Runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.

[0148] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0149] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0150] System libraries can include multiple functional modules. For example: surface manager, 2D graphics engine (e.g., SGL), 3D graphics processing library (e.g., OpenGL ES), media libraries, etc.

[0151] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0152] A 2D graphics engine is a graphics engine for 2D drawing.

[0153] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0154] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0155] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, audio driver, Wi-Fi driver, sensor driver, and Bluetooth driver.

[0156] It should be understood that the components included in the software structure shown in Figure 3 do not constitute a specific limitation on the electronic device 200. In other embodiments of this application, the electronic device 200 may include more or fewer components than shown, or combine some components, or split some components, or have different component arrangements.

[0157] Based on the aforementioned viewing system, this application provides a video interaction method. The technical solution of this application will be described in conjunction with the embodiments below.

[0158] Since this application involves the application of deep learning, it will be explained first for ease of understanding.

[0159] (1) Neural Network

[0160] Neural Networks (NNs) are machine learning models. A neural network can be composed of neural units, which are computational units that take xs and an intercept of 1 as input. The output of such a computational unit can be:

[0161] Where s = 1, 2, ..., n, n is a natural number greater than 1, W s For x sThe weights are denoted by b, and the bias of the neural unit is denoted by f(). The activation functions of the neural unit introduce non-linear characteristics into the neural network, converting the input signal into the output signal. The output signal of this activation function can be used as the input to the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of these individual neural units together; that is, the output of one neural unit can be the input of another. The input of each neural unit can be connected to the local receptive field of the previous layer to extract features from the local receptive field, which can be a region composed of several neural units.

[0162] (2) Deep Neural Networks

[0163] Deep neural networks (DNNs), also known as multilayer neural networks, can be understood as neural networks with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three layers based on their position: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs appear complex, the operation of each layer is actually quite simple, resembling a linear relationship as follows: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained through such a simple operation. Because DNNs have many layers, the coefficients W and the offset vector... The number of these parameters is therefore quite large. The definitions of these parameters in a DNN are as follows: Taking the coefficient W as an example: Assuming a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W resides, while the subscript corresponds to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the k-th neuron in layer L-1 to the j-th neuron in layer L are defined as follows: It's important to note that the input layer does not have a W parameter. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can perform more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrix of all layers in the trained deep neural network (a weight matrix formed by the vectors W from many layers).

[0164] (3) Convolutional Neural Network

[0165] A convolutional neural network (CNN) is a deep neural network with convolutional structures. It is a deep learning architecture, which refers to learning at multiple levels of abstraction using machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, where each neuron responds to an input image. A CNN contains a feature extractor consisting of convolutional layers and pooling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution with a trainable filter and an input image or a convolutional feature map.

[0166] A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution processing on the input signal. A convolutional layer can contain multiple convolution operators, also called kernels. In image processing, these operators act as filters, extracting specific information from the input image matrix. Essentially, a convolution operator can be a weight matrix, which is usually predefined. During the convolution operation, the weight matrix typically processes the input image pixel by pixel (or two pixels by two pixels, depending on the stride) along the horizontal direction, thus extracting specific features from the image. The size of the weight matrix should be related to the image size. It's important to note that the depth dimension of the weight matrix is ​​the same as the depth dimension of the input image; during convolution, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a single-depth convolutional output. However, in most cases, multiple weight matrices of the same size (rows × columns) are used instead of a single weight matrix. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. This dimension can be understood as being determined by the "multiple" factors mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix can be used to extract edge information, another to extract specific colors, and yet another to blur unwanted noise. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these weight matrices also have the same size. These extracted feature maps are then merged to form the output of the convolution operation. The weight values ​​in these weight matrices need to be obtained through extensive training in practical applications. The weight matrices formed by these trained weight values ​​can be used to extract information from the input image, enabling the convolutional neural network to make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layers often extract more general features, which can also be called low-level features. As the depth of the convolutional neural network increases, the features extracted by later convolutional layers become increasingly complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem being solved.

[0167] Because it's often necessary to reduce the number of training parameters, pooling layers are frequently introduced periodically after convolutional layers. This can be a single convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. Average pooling calculates the average value of pixel values ​​within a specific range as the result of average pooling. Max pooling takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.

[0168] After processing by convolutional / pooling layers, a convolutional neural network (CNN) is still insufficient to output the required information. As mentioned earlier, convolutional / pooling layers only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the CNN needs to utilize neural network layers to generate one or a set of desired class numbers of output. Therefore, the neural network can include multiple hidden layers, the parameters of which can be pre-trained based on training data relevant to a specific task type, such as image recognition, image classification, image super-resolution reconstruction, etc.

[0169] Optionally, after the multiple hidden layers in the neural network, there is also an output layer of the entire convolutional neural network. This output layer has a loss function similar to the classification cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the backpropagation will begin to update the weight values ​​and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.

[0170] (4) Recurrent Neural Network

[0171] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, the layers from the input layer to the hidden layer and then to the output layer are fully connected, but the nodes within each layer are unconnected. While this type of neural network has solved many difficult problems, it remains inadequate for many others. For example, predicting the next word in a sentence generally requires using the preceding words because words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is related to the outputs of previous sequences. Specifically, the network memorizes previous information and applies it to the calculation of the current output; that is, nodes within the same hidden layer are no longer unconnected but connected, and the input to a hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time step. Theoretically, RNNs can process sequential data of any length. Training an RNN is similar to training a traditional CNN or DNN. This algorithm also uses the backpropagation algorithm, but with one key difference: when an RNN is expanded, its parameters, such as W, are shared; however, this is not the case with traditional neural networks as illustrated above. Furthermore, in gradient descent, the output at each step depends not only on the network at the current step but also on the states of the network in previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).

[0172] Since we already have convolutional neural networks (CNNs), why do we need recurrent neural networks (RNNs)? The reason is simple. CNNs rely on the fundamental assumption that elements are independent of each other, and that input and output are also independent—like a cat and a dog. However, in the real world, many elements are interconnected. For example, stock prices fluctuate over time. Or, imagine someone saying, "I love traveling, and my favorite place is Yunnan. I definitely want to go there someday." Humans know the answer to this question is "Yunnan." Humans can infer from context, but how can machines do the same? This is where RNNs come in. RNNs aim to give machines the ability to remember, just like humans. Therefore, the output of an RNN depends on both the current input information and historical memory information.

[0173] (5) Loss Function

[0174] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value. Based on the difference, we update the weight vector of each layer (usually pre-configuring parameters before the initial update). For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network predicts the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss.

[0175] (6) Backpropagation algorithm

[0176] Convolutional neural networks can employ backpropagation (BP) to correct the parameters in the initial super-resolution model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates an error loss; this error loss information is then propagated back to update the parameters in the initial super-resolution model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the super-resolution model, such as the weight matrix.

[0177] (7) Generative Adversarial Networks

[0178] Generative adversarial networks (GANs) are a type of deep learning model. This model comprises at least two modules: a generative model and a discriminative model. These two modules learn from each other through a game-like interaction, resulting in better outputs. Both the generative and discriminative models can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GANs is as follows: Taking an image-generating GAN as an example, suppose there are two networks, G (Generator) and D (Discriminator). G is a network that generates images by receiving random noise z and using this noise, denoted as G(z). D is a discriminative network used to determine whether an image is "real." Its input parameter is x, representing an image, and its output D(x) represents the probability that x is a real image. A value of 1 indicates that the image is 100% real, while a value of 0 indicates that the image is impossible to be real. During the training of this generative adversarial network (GAN), the goal of the generative network G is to generate realistic images to deceive the discriminator network D, while the goal of the discriminator network D is to distinguish the images generated by G from real images as much as possible. Thus, G and D constitute a dynamic "game," which is the "adversarial" aspect of the GAN. Ideally, the game will result in G generating images G(z) that are sufficiently realistic, while D struggles to determine whether the images generated by G are real or not, i.e., D(G(z)) = 0.5. This yields a superior generative model G that can be used to generate images.

[0179] Figure 4 is a flowchart of process 400 of the video interaction method provided in this application. Process 400 can be executed by the aforementioned electronic device with the assistance of a video platform. The executing entity can be an application (APP) installed on the electronic device, such as a video APP, or it can be a video-related system application of the electronic device, or it can be a video-related plugin provided by the operating system of the electronic device, etc., without specific limitations. Process 400 is described as a series of steps or operations. It should be understood that process 400 can be executed in various orders and / or occur simultaneously, and is not limited to the execution order shown in Figure 4. Process 400 may include:

[0180] Step 401: In the story interaction mode, when the first video is played to a story interaction frame, the first video is paused and a first question is displayed on the playback interface. This first question is related to the story expressed in the story interaction frame.

[0181] The narrative interaction mode is a video interaction method provided in this application during video playback. After determining to enter narrative interaction mode, when the first video (e.g., a film, short video, etc.) reaches a narrative interaction frame, the user can interact with the plot-related information through a narrative Q&A format (including questions and answers). The user can select a question they want to know (the first question), which can be related to the plot expressed in the narrative interaction frame, such as the character's personality or motivation in the frame, and then obtain a corresponding answer (the first answer). This allows the user to understand the plot promptly, capture the information the creators of the first video want to convey, and thus better comprehend the story.

[0182] Before step 401, the method for determining whether to enter the story interaction mode in this application may include: before determining whether to enter the story interaction mode, when playing the scheduled interaction frame of the first video, displaying a prompt message, which prompts the user to confirm whether to enter the story interaction mode, the scheduled interaction frame being n seconds earlier than the story interaction frame, where n≥1. There are two possible branches after this:

[0183] 1) Upon receiving a fourth command triggered by the user, the system responds to the fourth command by determining whether to enter the story interaction mode. This fourth command includes instructions generated by the user operating the remote control or clicking a control on the touchscreen to indicate confirmation.

[0184] 2) Upon receiving the fifth instruction triggered by the user, in response to the fifth instruction, determine not to enter the story interaction mode and continue playing the first video. This fifth instruction includes instructions generated by the user operating the remote control or clicking a control on the touchscreen to indicate negation.

[0185] A scheduled interaction frame is pre-set in the first video. This scheduled interaction frame is earlier than the story interaction frame. For example, the scheduled interaction frame is 5 seconds earlier than the story interaction frame. In a 24-frame video, the scheduled interaction frame is about 120 frames earlier than the story interaction frame.

[0186] When the scheduled interactive frame is played, a prompt message can be displayed on the playback interface to ask the user to confirm whether to enter the interactive story mode. For example, a bubble-like prompt might appear on the playback control bar, displaying the message: "High-energy content ahead! Press the up button to enter interactive story mode, press the down button to not enter interactive story mode." The up or down button interaction can be tailored to a TV or set-top box scenario. Clicking the up button on the TV or set-top box remote triggers the process of entering interactive story mode, while clicking the down button prevents entry. Therefore, the fourth instruction can be generated by the user clicking the up button on the remote, and the fifth instruction can be generated by the user clicking the down button. Other interaction methods can also be used, such as confirmation or rejection controls on a touchscreen, allowing the user to select whether to enter or not enter interactive story mode. This application does not specifically limit the interaction method.

[0187] Optionally, if no user-triggered instruction is received after a preset waiting period, the system will not enter the story interaction mode and will continue playing the first video.

[0188] If the user neither clicks the confirmation control nor the rejection control when the appointment interaction frame is played, it can be determined that the story interaction mode will not be entered after waiting for a certain period of time (this period is preset, for example, 2 seconds).

[0189] After the process of determining whether to enter the story interaction mode, if it is determined not to enter the story interaction mode, the first video will continue to play and there will be no interaction with the user at the story interaction frame; if it is determined to enter the story interaction mode, step 401 will be executed.

[0190] As shown in Figure 5 (Figure 5 is a schematic diagram of the playback interface of this application), during the normal playback of the first video, the scheduled interaction frame is played first, and then the plot interaction frame is played. If the user selects to enter the plot interaction mode when playing the scheduled interaction frame, then a problem (Problem 1) is displayed on the playback interface (e.g., near the progress bar) when playing the plot interaction frame.

[0191] Interactive story frames serve as user interfaces. Questions and answers (including questions and answers) related to the storyline expressed in these frames can be obtained through offline preprocessing on the video platform, then transmitted to the electronic device via a communication link and saved. The process of the video platform generating questions and answers related to the storyline expressed in the interactive story frames offline may include:

[0192] 1) Obtain candidate frames of the first video, which are used to represent shots of the first video;

[0193] The first video is segmented into segments corresponding to several shots. The shot segmentation detection algorithm may include a threshold-based histogram analysis method, which compares the differences in color or brightness histograms of adjacent frames and considers a shot change to have occurred when the difference exceeds a preset threshold; or it may include a classifier or deep learning model judgment method, which automatically learns and identifies shot change patterns by training on a large amount of labeled data.

[0194] For any given video segment, the video frames contained therein are sampled and extracted. This sampling and extraction method can use a third-party library (such as FFmpeg) to sample the original video segment at equal intervals. The extracted video frames are then filtered out for motion blur using a sharpness detection algorithm such as the Laplacian operator gradient function. After deduplication of several consecutive frames, at least one of the most representative frames is retained as a candidate frame for that shot.

[0195] 2) Perform facial recognition on each candidate frame to obtain interactive plot frames, which include characters in the play facing the audience and occupying a central position.

[0196] Face detection and recognition, as well as face pose recognition, are performed on candidate frames to identify frames where the protagonist's face is clear and facing the audience directly, and then facial expression recognition is performed on them. Optionally, frames with dramatic tension (e.g., expressions of anger, joy, or fear) can be given higher weights and ranked higher. For example, taking the scene where Zhang San panics and mounts his horse to escape when his camp is attacked by a rival tribe, the following algorithm is used to rank candidate frames and obtain the dramatic interaction frames for that scene.

[0197] a) Face detection algorithms can include deep learning-based end-to-end real-time object detection algorithms such as You Only Look Once (YOLO), which can detect and locate multiple objects simultaneously. Because YOLO divides the entire image into a grid and predicts the object's class and bounding box on each grid, it is typically faster than other region-based object detection algorithms. YOLO, as mentioned above, is an object detection algorithm whose function is to detect the location and class of all objects in an image.

[0198] (b) Face recognition algorithms are mainly used to identify whether a detected face belongs to the main character in a film or television work. For example, it may involve first detecting faces in the input image to locate the face, then detecting key facial feature points from the face region, normalizing and aligning the face region using these key facial feature points, then extracting features from the normalized face region to obtain a face feature vector, and finally searching and matching the face feature vector with a pre-established database of main character face features, and determining the main character corresponding to the face feature vector based on the similarity score between the two.

[0199] c) The face pose recognition algorithm may include a third-party library (such as OpenCV) that identifies the frontal orientation of a face based on the haarcascade_frontalface_alt2.xml configuration file. OpenCV, as mentioned above, provides face pose detection capabilities, which can be used in this application to identify whether a face is frontal or profile, etc.

[0200] d) Facial expression recognition algorithms may include first detecting facial regions and extracting facial features from the input image, and then identifying specific emotions or facial expressions. The extracted facial features can be input into machine learning or deep learning models, such as Support Vector Machines (SVMs) and neural networks. After training, the model can map facial features to corresponding expression categories, such as happy, sad, and angry, thereby achieving facial expression recognition.

[0201] After the above processing, a candidate frame with a clear face and the person facing the camera can be selected from the candidate frames as the interactive candidate frame, which can be used as the plot interaction frame for that shot. Optionally, the frame number of the plot interaction frame for that shot can be persistently saved.

[0202] 3) Obtain plot-related information from the first video and generate a plot description text;

[0203] Based on the plot summaries and reviews of the aforementioned films and television works obtained from publicly available channels, the plot summary text can be acquired. Then, based on a natural language processing model and the plot summary text, preprocessing is performed to extract features from the text and store them in a feature library. Combined with the extracted plot interaction frames, prompts are generated. For example, the prompt might be: "The current plot involves Wang Wu attacking Zhang San's tribe's camp to avenge Zhao Liu's past theft of his wife. Li Si is forcibly abducted by Sun Qi. Please act as Zhang San in the first person, generating questions related to your own plot and providing corresponding answers." An example is shown below.

[0204] Question: Zhang San, why did you choose to abandon your wife and run away?

[0205] Answer: In order to survive and for the future revival of my family, I had to make the difficult decision to temporarily leave my wife, Li Si, to avoid danger. This escape was actually to preserve strength, seek allies, and ultimately return to unify the nation and begin my conquest. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of my people.

[0206] Character names are identified from the synopsis text of the first video using named entity recognition. Based on film and television metadata (e.g., cast list, including the association between actor names and character names), a mapping relationship is established between actors and character names. Then, using Optical Character Recognition (OCR) / Automatic Speech Recognition (ASR) technology, timestamped subtitles are extracted from the video file of the film and television work. Dialogue is matched with characters using speaker recognition technology, or the subtitle dialogue script file accompanying the film and television content (which already includes information on which character the dialogue belongs to) can be directly selected.

[0207] 4) Align the interactive story frames with the story description text to obtain at least one set of questions and answers (including questions and answers) for the interactive story frames.

[0208] Align video footage shots with the synopsis text of the film or television work. That is, align the segments from the previous steps with the synopsis text using feature extraction and clue matching. This can include the following methods:

[0209] Character identity matching: Establish a character identity recognition system to compare the characters mentioned in the plot introduction text with the characters appearing in the video, spread the influence of the characters through the shots within a certain time window, and calculate the similarity of the character identities.

[0210] Keyword matching: By comparing sentences in the plot summary text with keywords (such as names, locations, items, or events) in the subtitles, matching points are found, and stop word filtering is performed to reduce noise.

[0211] The following alignment algorithm can be used to align the image with the text modality:

[0212] Similarity calculation: Combining the character identity matching score and the subtitle keyword matching score, a similarity score is calculated for each sentence and shot.

[0213] DTW Algorithm: This algorithm solves the alignment problem using dynamic programming (DTW3). It considers the non-linear matching between shots and sentences, allowing for a certain range of time offsets while controlling the number of shots allocated to sentences to avoid overcrowding of a single sentence.

[0214] Based on the above process, at least one set of question-and-answer text can be obtained from the plot introduction text. This text is saved so that when the plot interaction frame is played, the corresponding set of question-and-answer text can be extracted and displayed.

[0215] Furthermore, based on the offline preprocessing process of the aforementioned video platforms, it is possible to add text-based dialogue questions and answers, using the characters' images and voiceprints to generate videos, which can then be used as video material in the interactive story mode.

[0216] 1) Creating personalized voice models: Using the authorized voice samples of the actors playing the characters in the play (e.g., avatar encoded as 001) as training samples and digital avatars as a basis, a voice model exclusive to the character is generated through speech synthesis.

[0217] 2) The generated question-and-answer text is fed into the model for inference, which can generate audio synthesized from the actors' personalized voices, such as question.wav and answer.wav in the following examples.

[0218] "quesion": "Zhang San, why did you choose to abandon your wife and run away?"

[0219] "Answer": "For survival and the future revitalization of the family, I had to make the difficult decision to temporarily leave my wife, Li Si, to avoid danger. This escape was actually to preserve strength, seek allies, and ultimately return to unify the nation and begin my journey. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of the people."

[0220] 3) Use AI technology to generate lip movement sequence frames that match the speech in question.wav and answer.wav, then align them to the face area to generate a video with questions and answers dictated by the characters in the show.

[0221] The speech synthesis module primarily ensures that the lip movements of the person speaking in the video are consistent with the audio content. It adjusts the lip movements based on the speech to synchronize the generated video character's lip movements with the input speech. This can include the following steps:

[0222] a) Extracting audio features: This is accomplished by using audio processing techniques such as spectrograms.

[0223] b) Extract video frames: Extract a series of consecutive video frames from the target video to be used as the target for lip animation.

[0224] c) Predicting lip movements: Using deep learning models, such as convolutional neural networks or recurrent neural networks, the correspondence between audio and lip movements is learned to generate lip animations suitable for the input audio.

[0225] d) Synthetic lip animation: The predicted lip motion sequence is applied to the lip region alignment and fusion of the target video.

[0226] e) Rendering and output: Combine the synthesized lip animation sequence with the content of the target video, and finally overlay the synthesized lip animation onto the target video for post-processing and adjustment.

[0227] f) When generating videos, keep both the start and end frames as interactive story frames to avoid abrupt transitions when switching between different videos.

[0228] The final generated video files, such as question.mp4 and answer.mp4, are combined with the previously generated question.wav and answer.wav using ffmpeg to create a synchronized audio-visual interactive video. This video is then transcoded into internet streaming media formats such as dash or hls and distributed via the origin server and CDN. The generated metadata is persistently stored on the video platform. This process yields at least one set of questions and answers from the synopsis text.

[0229] In the interactive story mode, when the playback reaches a story interaction frame of the first video, the first video is paused, and the first question is displayed on the playback interface. At this time, the background of the playback interface can be the story interaction frame. Based on the aforementioned offline generation process of questions and answers related to the story expressed by the story interaction frame, the first question can be displayed in the following two ways:

[0230] (1) Display the first question in text form on the playback interface.

[0231] The text for the first question comes from at least one set of question-and-answer text obtained from the offline preprocessing process of the aforementioned video platform. One question per interactive plot frame can be displayed sequentially on the playback interface, or multiple questions corresponding to interactive plot frames can be displayed simultaneously; there is no specific limitation on this. Optionally, the interactive plot frame can be used as the background at this time.

[0232] When the playback interface displays the text of a question, the first question mentioned above is that question; or, when the playback interface displays the text of multiple questions, the first question can be one of the multiple questions. The user can use the remote control or touch screen to tap the text of the first question to generate a first instruction, which indicates the question selected by the user (i.e., the first question) or indicates that the user chooses to answer the currently displayed question (the first question).

[0233] (2) Play the video of the first question on the playback interface. The video of the first question includes the target character uttering the first question. The target character is a character in the play who faces the audience and occupies a central position in the interactive plot frame. The video of the first question comes from at least one set of question-and-answer videos obtained from the offline preprocessing process of the aforementioned video platform. The video of one question (the first question) in the interactive plot frame can be played one at a time on the playback interface in a turn-by-turn manner. The user can operate the remote control or touch screen to click the control indicating confirmation in the video of the first question to generate a first instruction. The first instruction instructs the user to select to answer the currently played question (the first question). Optionally, the aforementioned target character can be presented as a digital human. The digital human image of the target character can be generated using 3D image technology, without specific limitations.

[0234] Step 402: After receiving the first instruction triggered by the user corresponding to the first question, in response to the first instruction, display the first answer on the playback interface. The first answer is used to answer the first question to help the user understand the plot expressed by the plot interaction frame.

[0235] The first answer may include at least one of the plot summary or film review of the first video, or the first answer may be based on at least one of the plot summary or film review of the first video, so that the user can understand the plot of the first video through the first answer, including the plot expressed by the plot interaction frame and the frame before it.

[0236] Corresponding to the first question above, the first answer can also be displayed in the following two ways:

[0237] (1) Display the first answer in text form on the playback interface. The text of the first answer comes from at least one set of question and answer text obtained from the offline preprocessing process of the aforementioned video platform. When the user selects the first question or chooses to answer the currently displayed first question, the text of the corresponding answer (the first answer) can be extracted and displayed on the playback interface.

[0238] (2) Play the video of the first answer on the playback interface. The video of the first answer includes the target character's oral first answer. The target character is a character in the play who faces the audience and occupies a central position in the interactive plot frame. The video of the first answer comes from at least one set of question-and-answer videos obtained from the offline preprocessing process of the aforementioned video platform. When the user selects to answer the first question currently being played, the corresponding answer (the first answer) video can be extracted and played on the playback interface.

[0239] In one possible implementation, after reading the first answer—whether the text or the video—the user can choose to view the next question (the second question) and its answer (the second answer). That is,

[0240] The system receives a second instruction triggered by the user, corresponding to a second question. This second question is related to the plot expressed in the interactive plot frame and is different from the first question mentioned above. In response to the second instruction, the system displays the second question on the playback interface. The aforementioned second instruction includes an instruction generated by the user operating a remote control or touchscreen to click on the text of the second question; or, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to click on a control indicating the next question.

[0241] In one possible implementation, after viewing the Nth answer (either the text or the video), the user can choose to stop watching the Q&A segments of the interactive storyline and continue watching the first video, where N ≥ 1.

[0242] The system receives a third instruction triggered by the user; in response to the third instruction, it continues playing the first video. The aforementioned third instruction includes instructions generated by the user operating a remote control or clicking a control on the touchscreen to indicate a return. After the user chooses not to continue watching the Q&A in the interactive storyline frame, playback of the first video can resume from the frame (interactive storyline frame) where the original video (first video) was interrupted. This does not exit the interactive storyline mode of the first video; rather, it stops displaying the undisplayed Q&A in the interactive storyline frame. That is, after watching the Nth answer, even if there are still unwatched Q&A in that interactive storyline frame, the user can still trigger the third instruction, and based on this third instruction, the undisplayed Q&A in the interactive storyline frame can be stopped, and playback of the first video can continue.

[0243] Optionally, if all questions and answers in the story interaction frame have been displayed after the Nth answer has been shown, the story interaction mode of that story interaction frame can be automatically exited and the first video can continue to play. In this case, no user-triggered command is required.

[0244] In one possible implementation, after viewing the Nth answer (either the text or the video), the user can choose to exit the interactive story mode and continue watching the first video. Even if a new interactive story frame is encountered during subsequent playback, the question and answer for that frame will not be displayed again, where N≥1. That is,

[0245] The system receives a sixth command triggered by the user; in response to the sixth command, it exits the story interaction mode and resumes playing the first video. The aforementioned sixth command includes commands generated by the user operating the remote control or tapping the exit control on the touchscreen. After the user selects to exit the story interaction mode, playback of the first video can resume from the frame (story interaction frame) where the original video (first video) was interrupted. At this point, the story interaction mode for the first video has exited; even if there is another story interaction frame, the questions and answers for the next story interaction frame will not be displayed when playback reaches that frame; instead, the first video will continue playing.

[0246] It is clear that the interactive story mode does not change the first video itself. Users can still watch the first video in its entirety without experiencing the confusion that can occur due to too many story branches, as is the case with related technologies.

[0247] This embodiment implements a story-driven interactive mode. After confirming the entry into this mode, when the first video plays a story-driven interactive frame, the user can select the questions they want to know and receive corresponding answers. Through this story-related Q&A, users can quickly understand the plot's direction, the characters' personalities and motivations, and detailed analyses, without needing to search social media interactions, read professional film reviews, or access external websites and forums. This allows users to better understand the information the creators want to convey, thus enhancing their viewing experience.

[0248] This application generates plot-related Q&A (including questions and answers) in the video and the plot interaction frames. When the plot interaction frames are played, the Q&A format allows the user to interact with the plot, making the audience more immersive in the plot and the characters' situations. This not only allows them to understand the plot in a timely manner, but also allows them to capture the information that the creators of the film and television works want to convey, thereby better understanding the plot and improving the user's viewing experience.

[0249] The technical solutions of the method embodiments shown in Figure 4 will be described in detail below using several specific examples.

[0250] Figure 6 is a flowchart illustrating the video interaction method of this application. While watching a film or television work, the user can enter a story interaction mode and learn about the plot through character-based question-and-answer text. This embodiment uses a film or television work featuring Zhang San as the protagonist as an example.

[0251] 1. The offline preprocessing process of the video platform includes the following steps:

[0252] 1) Perform shot segmentation processing on the video files of the aforementioned film and television works, dividing the entire video into video segments corresponding to several shots. The shot segmentation detection algorithm may include a threshold-based histogram analysis method, that is, by comparing the differences in color or brightness histograms of adjacent frames, a shot change is considered to have occurred when the difference exceeds a preset threshold; or, it may include a classifier or deep learning model judgment method, that is, by training a large amount of labeled data, the shot change pattern is automatically learned and identified.

[0253] 2) For any given video segment, sample and extract its constituent video frames. This sampling and extraction method can use a third-party library (such as FFmpeg) to sample the original video segment at equal intervals. The extracted video frames are then filtered to remove motion-blurred frames using a sharpness detection algorithm, such as the Laplacian operator gradient function. After deduplication of several consecutive frames, at least one of the most representative frames is retained as a candidate frame for that shot. FFmpeg, as mentioned above, can be used to record and convert digital audio and video, converting them to streaming formats, and performing encoding and decoding operations on various platforms. FFmpeg is an open-source, cross-platform multimedia framework created by Fabrice Bellard, primarily used for processing audio and video data. It provides a set of libraries and tools for processing audio and video data, enabling operations such as audio and video format conversion, editing, trimming, merging, and playback. FFmpeg supports many audio and video formats, and due to its powerful features and wide application, it can be considered a standard in the field of audio and video processing. The Laplacian operator is a second-order differential operator primarily used in image processing and computer vision. Its functions include: (1) Edge detection: The Laplacian operator can detect edges in an image because it can highlight the high-frequency parts of the image, i.e., edges; (2) Image enhancement: By performing a Laplacian transform on an image, the contrast and clarity of the image can be enhanced; (3) Feature extraction: The Laplacian operator can be used to extract features in an image, such as texture and shape; (4) Noise removal: By performing a Laplacian transform on an image, noise in the image can be removed; (5) Image segmentation: The Laplacian operator can be used for image segmentation, such as separating the foreground and background in an image. It can be seen that the Laplacian operator has a wide range of applications in the fields of image processing and computer vision, and can be used for various image processing tasks.

[0254] 3) Perform face detection and recognition, as well as face pose recognition, on the candidate frames to identify frames where the protagonist's face is clear and facing the audience, and then perform expression recognition on them. Optionally, frames with dramatic tension (e.g., expressions of anger, joy, or fear) can be given higher weights and ranked higher. For example, taking the scene where Zhang San panics and mounts his horse to escape when his camp is attacked by a hostile tribe, the following algorithm can be used to rank the candidate frames and obtain the interactive frames of that scene.

[0255] a) Face detection algorithms can include deep learning-based end-to-end real-time object detection algorithms such as You Only Look Once (YOLO), which can detect and locate multiple objects simultaneously. Because YOLO divides the entire image into a grid and predicts the object's class and bounding box on each grid, it is typically faster than other region-based object detection algorithms. YOLO, as mentioned above, is an object detection algorithm whose function is to detect the location and class of all objects in an image.

[0256] (b) Face recognition algorithms are mainly used to identify whether a detected face belongs to the main character in a film or television work. For example, it may involve first detecting faces in the input image to locate the face, then detecting key facial feature points from the face region, normalizing and aligning the face region using these key facial feature points, then extracting features from the normalized face region to obtain a face feature vector, and finally searching and matching the face feature vector with a pre-established database of main character face features, and determining the main character corresponding to the face feature vector based on the similarity score between the two.

[0257] c) The face pose recognition algorithm may include a third-party library (such as OpenCV) that identifies the frontal orientation of a face based on the haarcascade_frontalface_alt2.xml configuration file. OpenCV, as mentioned above, provides face pose detection capabilities, which can be used in this application to identify whether a face is frontal or profile, etc.

[0258] d) Facial expression recognition algorithms may involve first detecting facial regions and extracting facial features from the input image, and then identifying specific emotions or facial expressions. The extracted facial features can be input into machine learning or deep learning models, such as Support Vector Machines (SVMs) and neural networks. After training, the model can map facial features to corresponding expression categories, such as happy, sad, and angry, thereby achieving facial expression recognition. The aforementioned SVM is a binary classification model, based on the principle of structural risk minimization in statistical learning theory. The main idea of ​​SVM is to partition the data using a hyperplane, ensuring that data points of different categories are correctly classified and maximizing the distance between the classification boundary and the nearest sample point, i.e., maximizing the classifier margin. The core of SVM is finding an optimal hyperplane that maximizes the distance from data points to this hyperplane. This optimal hyperplane is called the "maximum margin hyperplane." In SVM, the data points closest to the hyperplane are called "support vectors," which play a crucial role in determining the hyperplane's position and orientation. The optimization problem of SVM can be transformed into a convex quadratic programming problem, and the maximum margin hyperplane can be obtained by solving this problem. In practical applications, since data may not be linearly separable, kernel functions are needed to map the data to a high-dimensional space, making the data linearly separable in that space. Therefore, SVM is a powerful classification algorithm with excellent generalization performance and high accuracy, and is widely used in many practical applications.

[0259] After the above processing, facial expressions can be weighted according to rules (facial expressions usually convey joy, anger, sorrow, and happiness; these expressions often have more dramatic tension than expressionless faces, so certain expressions can be given higher priority according to business needs). Then, a candidate frame with a clear face and the character facing the camera is selected from the candidate frames as the interaction candidate frame, which can be used as the drama interaction frame for that shot. Optionally, the frame number of the drama interaction frame for that shot can be persistently saved. For example, the saved shot metadata is as follows, where the shot_id field is the shot number, the keyframe_id field is the frame number of the drama interaction frame, and the keyframe_avatar is the character code in the drama:

[0260] 4) Based on the plot summaries and reviews of the aforementioned films and television works obtained from publicly available channels, the plot summary text can be acquired. Then, based on a natural language processing model and the plot summary text, preprocessing is performed to extract features from the text and store them in a feature library. Combined with the extracted plot interaction frames, prompts are generated. For example, the prompt might be: "The current plot involves Wang Wu attacking Zhang San's tribe's camp to avenge Zhao Liu's past theft of his wife. Li Si is forcibly abducted by Sun Qi. Please play the role of Zhang San in the first person, generating questions related to your own plot and providing corresponding answers." An example is shown below:

[0261] Question: Zhang San, why did you choose to abandon your wife and run away?

[0262] Answer: In order to survive and for the future revival of my family, I had to make the difficult decision to temporarily leave my wife, Li Si, to avoid danger. This escape was actually to preserve strength, seek allies, and ultimately return to unify the nation and begin my conquest. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of my people.

[0263] 5) Identify character names from the synopsis text of film and television works using named entity recognition, and establish a mapping relationship between actors and character names based on film and television metadata (e.g., cast list, including the association between actor names and character names). Extract timestamped subtitles from the video files of the film and television works using Optical Character Recognition (OCR) / Automatic Speech Recognition (ASR) technology, and match dialogue with characters using speaker recognition technology, or directly use the subtitle dialogue script file that comes with the film and television content (which already contains information about which character the dialogue belongs to). The aforementioned named entity recognition algorithm includes the following recognition process: (1) Word segmentation: dividing the text into words or word sequences, which is a preliminary step for named entity recognition; (2) Feature extraction: extracting features from the text, such as part of speech, context, word roots, etc., to determine whether a word is a named entity; (3) Data annotation: using existing annotation datasets to annotate the named entities in the text, which can be used to train the model; (4) Model training: using the annotation dataset to train the model, commonly used models include Conditional Random Field (CRF), Maximum Entropy Model (MaxEnt), Support Vector Machine (SVM), etc.; (5) Prediction: using the trained model to perform named entity recognition on new text and output the recognition results; (6) Post-processing: performing post-processing on the recognition results, such as merging adjacent named entities and removing erroneous recognition results.

[0264] 6) Align the video footage with the synopsis text of the film / TV show. That is, align the segments from the previous steps with the synopsis text using feature extraction and clue matching. This can include the following methods:

[0265] Character identity matching: Establish a character identity recognition system to compare the characters mentioned in the plot introduction text with the characters appearing in the video, spread the influence of the characters through the shots within a certain time window, and calculate the similarity of the character identities.

[0266] Keyword matching: By comparing sentences in the plot summary text with keywords (such as names, locations, items, or events) in the subtitles, matching points are found, and stop word filtering is performed to reduce noise.

[0267] The following alignment algorithm can be used to align the image with the text modality:

[0268] Similarity calculation: Combining the character identity matching score and the subtitle keyword matching score, a similarity score is calculated for each sentence and shot.

[0269] DTW Algorithm: This algorithm solves the alignment problem using dynamic programming (DTW3). It considers the non-linear matching between shots and sentences, allowing for a certain range of time offsets while controlling the number of shots allocated to sentences to avoid overcrowding of a single sentence.

[0270] For example, the synopsis text corresponding to shot 265 is "Wang Wu, seeking revenge for Zhao Liu's past act of stealing his wife, raids Zhang San's tribe's camp; Li Si is forcibly abducted by Sun Qi." This shot is represented by the interactive frame (frame 14363), where the avatar is the character's code, corresponding to multiple question-and-answer pairs. This establishes a connection between unstructured video and synopsis text through the shots, forming structured data corresponding to the video footage and synopsis text, which is then persistently saved to the video platform. An example is shown below:

[0271] 2. The client provides a real-time interactive storyline experience, as shown in Figure 6, including the following steps:

[0272] 01. The story interaction mode relies on the question-and-answer related data structure corresponding to the story introduction text obtained from the backend interface, as shown in the following example:

[0273] The core information in the JSON example above is that the current video segment contains several plot interaction frames. Each plot interaction frame has at least one related question and an answer from the perspective of a plot character. This structured data is parsed and cached in memory for subsequent plot interactions.

[0274] To improve user experience, background data retrieval can begin before or at the very start of playback. This allows for faster display when the user selects to enter the interactive story mode, reducing loading delays. Real-time retrieval will incur additional waiting time.

[0275] The client can determine whether it is approaching an interactive storyline frame based on the playback progress frame number. When it approaches frame 14243 (approximately 5 seconds before frame 14363 in the example, or about 120 frames ahead in a 24-frame video), a bubble-like prompt appears on the playback control bar, displaying the message: "High-energy content ahead! Press the up button to enter interactive storyline mode." The up button interaction can be for TV or set-top box scenarios, or for touchscreens, it can be triggered by clicking buttons such as "View Question," "Next Question," or "Back." If the user does not trigger an interaction, the prompt will disappear automatically after frame 14363. The end-user experience is that when watching a video with interactive storyline features, the playback control bar will provide a prompt when the playback progress approaches the nth interactive frame, allowing the user to decide whether to enter interactive storyline mode. If the user ignores the prompt, playback continues.

[0276] 02. After the user confirms entry into the interactive story mode, the playback window remains on the interactive story frame. This frame is a shot of the character Zhang San in the story, based on content understanding pre-analysis. The client uses this frame as the background for the interactive story mode, loads pre-processed questions from memory, and displays question 1 at the bottom of the screen, for example:

[0277] "Zhang San, why did you choose to abandon your wife and run away?"

[0278] Press the OK button to view the answer, press the right button to move on to the next question, and press the left button to exit the story interaction mode.

[0279] 03. When a user interacts to view the answer to question 1, the pre-generated answer text and other information about question 1 will be displayed. For example:

[0280] "In order to survive and for the future to revive the family, I had to make the difficult decision to temporarily leave my wife, Li Si, to avoid danger. This escape was actually to preserve strength, seek allies, and eventually return to unify the nation and begin my journey. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of the people."

[0281] 04. If the user is not very interested in the current question, they can switch to the next question by right-clicking. The question and answer text of question 2 can be loaded and displayed from the pre-stored structured data in memory.

[0282] 05. Users can decide whether to view the answer to question 2. If they do, the answer text and other information related to question 2 will be displayed.

[0283] 06. Users can continue to trigger the loading of new questions through interaction and view the corresponding answers, and so on.

[0284] 07. When users want to continue watching the video, they can exit the story interaction mode through interaction. At this time, the story interaction mode interface will be destroyed, the interface will be restored to the playback window, and playback will continue from frame n of the story interaction.

[0285] This embodiment generates dialogue text related to characters in film and television works. When the corresponding interactive frame is played, the interactive mode is triggered to display the dialogue text, making the audience more immersive in the plot and the characters' situations, easier to understand the plot, and enhancing the viewing experience. Furthermore, users can understand the plot with the help of rich multimedia information while watching, thus better empathizing with the characters in the film and television work.

[0286] Figure 7 is a flowchart illustrating the video interaction method of this application. While watching a film or television work, users can enter a story interaction mode and learn about the plot through character-based question-and-answer videos. This embodiment uses a film or television work featuring Zhang San as the protagonist as an example.

[0287] 1'. Based on the offline preprocessing process of the video platform in the embodiment shown in Figure 6, this embodiment adds the function of generating videos from text-based dialogue questions and answers using the characters' images and voiceprints, which are then used as video material in the interactive dialogue mode.

[0288] 1) Creating personalized voice models: Using the authorized voice samples of the actors playing the characters in the play (e.g., avatar encoded as 001) as training samples and digital avatars as a basis, a voice model exclusive to the character is generated through speech synthesis.

[0289] 2) The generated question-and-answer text is fed into the model for inference, which can generate audio synthesized from the actors' personalized voices, such as question.wav and answer.wav in the following examples.

[0290] "quesion": "Zhang San, why did you choose to abandon your wife and run away?"

[0291] "Answer": "For survival and the future revitalization of the family, I had to make the difficult decision to temporarily leave my wife, Li Si, to avoid danger. This escape was actually to preserve strength, seek allies, and ultimately return to unify the nation and begin my journey. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of the people."

[0292] 3) Use AI technology to generate lip movement sequence frames that match the speech in question.wav and answer.wav, then align them to the face area to generate a video with questions and answers dictated by the characters in the show.

[0293] The speech synthesis module primarily ensures that the lip movements of the person speaking in the video are consistent with the audio content. It adjusts the lip movements based on the speech to synchronize the generated video character's lip movements with the input speech. Relevant tools can include wav2lip, or utilize third-party closed-source voice and lip-syncing services such as Hygen. The process may include the following steps:

[0294] a) Extracting audio features: This is accomplished by using audio processing techniques such as spectrograms.

[0295] b) Extract video frames: Extract a series of consecutive video frames from the target video to be used as the target for lip animation.

[0296] c) Predicting lip movements: Using deep learning models, such as convolutional neural networks or recurrent neural networks, the correspondence between audio and lip movements is learned to generate lip animations suitable for the input audio.

[0297] d) Synthetic lip animation: The predicted lip motion sequence is applied to the lip region alignment and fusion of the target video.

[0298] e) Rendering and output: Combine the synthesized lip animation sequence with the content of the target video, and finally overlay the synthesized lip animation onto the target video for post-processing and adjustment.

[0299] f) When generating videos, keep both the start and end frames as interactive story frames to avoid abrupt transitions when switching between different videos.

[0300] The final generated video files, such as question.mp4 and answer.mp4, are combined with the previously generated question.wav and answer.wav using ffmpeg to create a synchronized audio-visual interactive video. This video is then transcoded into internet streaming media formats such as dash or hls and distributed via the origin server and CDN. The generated metadata is persistently stored on the video platform. For example, the saved scene metadata is shown below, where https: / / ip:port / path / question_122347fsdfw3rwer.mpd is the playable address of the interactive video.

[0301] The core information in the JSON example above is that the current video contains several interactive story frames. Each interactive story frame has at least one related question and a video link to the answer from the perspective of the story character. This structured data is parsed and cached in memory for subsequent interactive story events.

[0302] To improve user experience, background fetching can begin before or at the very start of playback. This allows for faster display when the user selects to enter the interactive story mode, reducing loading latency. Real-time fetching will incur additional waiting time. Additionally, to further accelerate story video loading, a player instance can be created in advance for quick video playback, reducing the waiting time for instance creation.

[0303] The client can determine whether it's approaching an interactive frame based on the playback progress frame number. When it's approximately 5 seconds before frame 14363 (about 120 frames ahead in a 24-frame video), i.e., at frame 14243, a bubble-like prompt appears on the playback control bar. The bubble's content could be, "High-energy content ahead! Press the up button to enter interactive mode." The up button interaction is for TV or set-top box scenarios; for touchscreens, it can be triggered by clicking. If the user doesn't trigger an interaction, the prompt disappears automatically after frame 14363. The end-user experience is that when watching a video with interactive features, the playback progress bar provides a prompt when the playback progress approaches the nth interactive frame, allowing the user to decide whether to enter interactive mode. If the user ignores the prompt, playback continues.

[0304] 02. Once the user confirms entry into the story interaction mode, a video corresponding to the question in that story interaction frame will play. For example, the video content might be a digital avatar of the character Zhang San saying, "Zhang San, why did you choose to abandon your wife and run away?" Below the video, there are prompts: Press the OK button to view the answer, press the right button to move on to the next question, and press the left button to exit the story interaction mode.

[0305] 03. When a user interacts to view the answer to Question 1, a pre-generated answer video begins to play. For example, a digital avatar of the character Zhang San says: "For survival and the future revitalization of my family, I had to make the difficult decision to temporarily leave my wife Li Si to avoid danger. This escape was actually to preserve strength, seek allies, and ultimately return to unify the nation and begin my journey. This separation, though filled with reluctance, was a necessary sacrifice in that turbulent era for a greater goal and the prosperity of my people."

[0306] 04. If the user is not very interested in the current question, they can switch to the next question's video by right-clicking. At this time, the player continues to load the next video playback address and starts playing the video. The resources occupied by the first video instance are released and reclaimed, and preparations are made for the next alternative video to play in the background.

[0307] 05. Users can trigger the playback of the answer video through interaction, such as pressing the OK button. The video playback address is, for example, https: / / ip:port / path / answer_1sdf3442sdfw3rwer.mpd as shown above.

[0308] 06. Users can continue to trigger the loading of new questions through interaction and view the corresponding answers, and so on.

[0309] 07. When users want to continue watching the video, they can exit the story interaction mode through interaction, such as pressing the left mouse button. At this time, the player in story interaction mode will be destroyed, the interface will return to the playback window, and playback will continue from frame n of the story interaction mode.

[0310] This embodiment generates dialogue videos related to characters in film and television works. When the corresponding interactive frame is played, the interactive mode is triggered to play the dialogue videos related to the characters. This allows viewers to feel more immersed in the plot and the characters' situations, making it easier to understand the story and enhancing the viewing experience. Moreover, users can understand the plot with the help of rich multimedia information while watching the film, thus better empathizing with the characters in the film and television works.

[0311] Figure 8 is a structural schematic diagram of the video interaction device 800 provided in this application. As shown in Figure 8, the video interaction device 800 of this embodiment can be applied to the aforementioned electronic device. The video interaction device 800 may include: a processing module 801 and a display module 802.

[0312] The display module 802 is configured to, after determining that the plot interaction mode is to be entered, pause the playback of the first video when the plot interaction frame of the first video is played, and display a first question on the playback interface, the first question being related to the plot expressed by the plot interaction frame; and, in response to receiving a first instruction triggered by the user corresponding to the first question, display a first answer on the playback interface, the first answer being used to answer the first question to help the user understand the plot expressed by the plot interaction frame.

[0313] In one possible implementation, the processing module 801 is configured to receive a second instruction triggered by a user corresponding to a second question, the second question being related to the plot expressed by the plot interaction frame and different from the first question; the display module 802 is further configured to display the second question on the playback interface in response to the second instruction.

[0314] In one possible implementation, the processing module 801 is further configured to receive a third instruction triggered by the user; the display module 802 is further configured to continue playing the first video in response to the third instruction.

[0315] In one possible implementation, the display module 802 is further configured to display a prompt message when playing the scheduled interaction frame of the first video before entering the story interaction mode. The prompt message is used to prompt the user to confirm whether to enter the story interaction mode. The scheduled interaction frame is n seconds earlier than the story interaction frame, where n≥1. The processing module 801 is further configured to determine whether to enter the story interaction mode in response to the fourth instruction triggered by the user after receiving the fourth instruction.

[0316] In one possible implementation, the processing module 801 is further configured to, upon receiving a fifth instruction triggered by the user, determine, in response to the fifth instruction, not to enter the story interaction mode and continue playing the first video.

[0317] In one possible implementation, the processing module 801 is further configured to, after waiting for a preset time, if no instruction triggered by the user is received, determine not to enter the story interaction mode and continue playing the first video.

[0318] In one possible implementation, the display module 802 is specifically used to display the first question in text form on the playback interface.

[0319] In one possible implementation, the first instruction includes instructions generated by the user operating a remote control or touchscreen to tap the text of the first question.

[0320] In one possible implementation, the display module 802 is specifically used to play a video of the first question on the playback interface. The video of the first question includes a target person uttering the first question. The target person is a character in the play who is facing the audience and occupies a central position in the interactive plot frame.

[0321] In one possible implementation, the first instruction includes an instruction generated by the user operating a remote control or touchscreen to click a control in the video of the first question that indicates confirmation.

[0322] In one possible implementation, the display module 802 is specifically used to display the first answer in text form on the playback interface; or, to play a video of the first answer on the playback interface, wherein the video of the first answer includes a target person dictating the first answer, and the target person is a character in the play who is facing the audience and occupies a central position in the interactive plot frame.

[0323] In one possible implementation, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to tap the text of the second question; or, the second instruction includes an instruction generated by the user operating a remote control or touchscreen to tap a control indicating the next question.

[0324] In one possible implementation, the third instruction includes instructions generated by the user operating a remote control or a touchscreen to tap a control indicating a return.

[0325] In one possible implementation, the fourth instruction includes instructions generated by the user operating a remote control or touchscreen to click on a control indicating confirmation.

[0326] In one possible implementation, the fifth instruction includes an instruction generated by the user operating a remote control or touchscreen to click a control that indicates negation.

[0327] In one possible implementation, the first answer includes at least one of the plot summary or film review of the first video; or, the first answer is obtained based on at least one of the plot summary or film review of the first video.

[0328] The apparatus in this embodiment can be used to execute the technical solution of the method embodiment shown in FIG4. Its implementation principle and technical effect are similar, and will not be described again here.

[0329] It is understood that, in order to achieve the above-mentioned functions, electronic devices include hardware and / or software modules that perform the respective functions. Based on the algorithmic steps of the examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in conjunction with the embodiments, but such implementation should not be considered beyond the scope of this application.

[0330] In one example, FIG9 shows a schematic block diagram of an apparatus 900 according to an embodiment of the present application. As shown in FIG9, the apparatus 900 may include a processor 901 and a transceiver / transceiver pin 902, and optionally, a memory 903.

[0331] The various components of device 900 are coupled together via bus 904, which includes a data bus, a power bus, a control bus, and a status signal bus. However, for clarity, all buses are referred to as bus 904 in the figure.

[0332] Optionally, the memory 903 can be used for the instructions in the foregoing method embodiments. The processor 901 can be used to execute the instructions in the memory 903, control the receive pin to receive signals, and control the transmit pin to transmit signals.

[0333] The device 900 may be an electronic device or a chip of an electronic device in the above method embodiments.

[0334] All relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0335] This embodiment also provides a computer storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned related method steps to implement the voice assistant-based control method in the above embodiment.

[0336] This embodiment also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the above-mentioned related steps to implement the voice assistant-based control method in the above embodiment.

[0337] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory. The memory is used to store computer execution instructions. When the apparatus is running, the processor can execute the computer execution instructions stored in the memory to cause the chip to execute the voice assistant-based control method in the above-described method embodiments.

[0338] In this embodiment, the electronic device, computer storage medium, computer program product or chip are all used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding method provided above, and will not be repeated here.

[0339] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0340] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0341] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0342] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0343] Any content from each embodiment of this application, as well as any content from the same embodiment, can be freely combined. Any combination of the above content is within the scope of the embodiments of this application.

[0344] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0345] The embodiments of this application have been described above with reference to the accompanying drawings. However, the embodiments of this application are not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the embodiments of this application without departing from the spirit of the embodiments of this application and the scope of protection of the claims, and all of these forms are within the protection scope of the embodiments of this application.

Claims

1. A video interaction method, characterized by, The method comprises: in a plot interaction mode, when a plot interaction frame of a first video is played, pausing playing of the first video, and displaying a first question on a playing interface, the first question being related to a plot expressed by the plot interaction frame; after receiving a first instruction triggered by a user and corresponding to the first question, in response to the first instruction, displaying a first answer on the playing interface, the first answer being used to answer the first question to help the user understand the plot expressed by the plot interaction frame.

2. The method of claim 1, wherein, after displaying the first answer on the playing interface, the method further comprises: receiving a second instruction triggered by the user and corresponding to a second question, the second question being related to the plot expressed by the plot interaction frame and being different from the first question; in response to the second instruction, displaying the second question on the playing interface.

3. The method according to claim 1 or 2, characterized in that, after displaying the first answer on the playing interface, the method further comprises: receiving a third instruction triggered by the user; in response to the third instruction, continuing playing of the first video.

4. The method according to any one of claims 1-3, characterized in that, The method further comprises: before entering the plot interaction mode, when a pre-arranged interaction frame of the first video is played, displaying prompt information, the prompt information being used to prompt the user to confirm whether to enter the plot interaction mode, the pre-arranged interaction frame being earlier than the plot interaction frame by n seconds, n≥1; after receiving a fourth instruction triggered by the user, in response to the fourth instruction, determining to enter the plot interaction mode.

5. The method of claim 4, wherein, after displaying the prompt information, the method further comprises: after receiving a fifth instruction triggered by the user, in response to the fifth instruction, determining not to enter the plot interaction mode, and continuing playing of the first video.

6. The method of claim 4, wherein, after displaying the prompt information, the method further comprises: after waiting for a preset time length, if no instruction triggered by the user is received, determining not to enter the plot interaction mode, and continuing playing of the first video.

7. The method according to any one of claims 1 to 6, characterized in that, The displaying of the first question on the playing interface comprises: displaying the first question in the form of text on the playing interface.

8. The method of claim 7, wherein, The first instruction comprises an instruction generated by the user operating a remote controller or a touch screen to click text of the first question.

9. The method according to any one of claims 1-6, characterized in that, The displaying of the first question on the playing interface comprises: playing a video of the first question on the playing interface, content of the video of the first question comprising the target character dictating the first question, the target character being a character in the plot who faces the audience and occupies a core position in the plot interaction frame.

10. The method of claim 9, wherein, The first instruction comprises an instruction generated by the user operating a remote controller or a touch screen to click a control representing confirmation in the video of the first question.

11. The method according to any one of claims 1-10, characterized in that, The displaying of the first answer on the playing interface comprises: displaying the first answer in the form of text on the playing interface; or playing a video of the first answer on the playing interface, content of the video of the first answer comprising the target character dictating the first answer, the target character being a character in the plot who faces the audience and occupies a core position in the plot interaction frame.

12. The method of claim 2, wherein, The second instruction includes an instruction generated by a user clicking a text of the second question by operating a remote controller or a touch screen; or the second instruction includes an instruction generated by a user clicking a control representing a next question by operating the remote controller or the touch screen.

13. The method of claim 3, wherein, The third instruction includes an instruction generated by a user clicking a control representing a return by operating the remote controller or the touch screen.

14. The method of claim 4, wherein, The fourth instruction includes an instruction generated by a user clicking a control representing a determination by operating the remote controller or the touch screen.

15. The method of claim 5, wherein, The fifth instruction includes an instruction generated by a user clicking a control representing a negation by operating the remote controller or the touch screen.

16. The method of any one of claims 1-15, wherein, Further comprising: obtaining a candidate frame of the first video, the candidate frame being used to represent a shot of the first video; performing face recognition on the candidate frame to obtain the plot interaction frame, the plot interaction frame including a character in the plot facing the audience and occupying a core position; obtaining plot-related information of the first video and generating a plot introduction text; aligning the plot interaction frame with the plot introduction text to obtain the first question and the first answer.

17. The method of any one of claims 1-16, wherein, The first answer includes at least one of a plot introduction or a film and television review of the first video; or The first answer is obtained based on at least one of a plot introduction or a film and television review of the first video.

18. A video interactive device, characterized by Comprising: a display module, configured to, after determining to enter a plot interaction mode, pause playing of a first video when a plot interaction frame of the first video is played, and display a first question on a playing interface, the first question being related to a plot represented by the plot interaction frame; after receiving a first instruction triggered by a user and corresponding to the first question, display a first answer on the playing interface in response to the first instruction, the first answer being used to answer the first question to help the user understand the plot represented by the plot interaction frame.

19. The apparatus of claim 18, wherein, Further comprising: a processing module, configured to receive a second instruction triggered by a user and corresponding to a second question, the second question being related to the plot represented by the plot interaction frame and being different from the first question; the display module is further configured to display the second question on the playing interface in response to the second instruction.

20. The apparatus of claim 18 or 19, wherein, the processing module is further configured to receive a third instruction triggered by the user; the display module is further configured to continue playing the first video in response to the third instruction.

21. The apparatus of any one of claims 18-20, wherein, the display module is further configured to, before entering the plot interaction mode, display prompt information when a pre-plot interaction frame of the first video is played, the prompt information being used to prompt the user to confirm whether to enter the plot interaction mode, the pre-plot interaction frame being n seconds earlier than the plot interaction frame, n≥1; the processing module is further configured to, after receiving a fourth instruction triggered by the user, determine to enter the plot interaction mode in response to the fourth instruction.

22. The apparatus of claim 21, wherein, the processing module is further configured to, after receiving a fifth instruction triggered by the user, determine not to enter the plot interaction mode in response to the fifth instruction, and continue playing the first video.

23. The apparatus of claim 21, wherein, The processing module is further configured to determine not to enter the interactive mode if no instruction triggered by the user is received after waiting for a preset time length, and continue playing the first video.

24. The apparatus of any one of claims 18-23, wherein, The display module is specifically configured to display the first question in the form of text on the playing interface.

25. The apparatus of claim 24, wherein, The first instruction includes an instruction generated by the user operating a remote controller or a touch screen to click the text of the first question.

26. The apparatus of any one of claims 18-23, wherein, The display module is specifically configured to play a video of the first question on the playing interface, and a content of the video of the first question includes the target character orally stating the first question. The target character is a character in the interactive frame who faces the audience and occupies a core position.

27. The apparatus of claim 26, wherein, The first instruction includes an instruction generated by the user operating a remote controller or a touch screen to click a control representing confirmation in the video of the first question.

28. The apparatus of any one of claims 18-27, wherein, The display module is specifically configured to display the first answer in the form of text on the playing interface; or The display module is specifically configured to play a video of the first answer on the playing interface, and a content of the video of the first answer includes the target character orally stating the first answer. The target character is a character in the interactive frame who faces the audience and occupies a core position.

29. The apparatus of claim 19, wherein, The second instruction includes an instruction generated by the user operating a remote controller or a touch screen to click the text of the second question; or the second instruction includes an instruction generated by the user operating a remote controller or a touch screen to click a control representing a next question.

30. The apparatus of claim 20, wherein, The third instruction includes an instruction generated by the user operating a remote controller or a touch screen to click a control representing return.

31. The apparatus of claim 21, wherein, The fourth instruction includes an instruction generated by the user operating a remote controller or a touch screen to click a control representing determination.

32. The apparatus of claim 22, wherein, The fifth instruction includes an instruction generated by the user operating a remote controller or a touch screen to click a control representing negation.

33. The apparatus of any one of claims 18-32, wherein, The first answer includes at least one of a plot introduction or a film and television review of the first video; or The first answer is obtained based on at least one of a plot introduction or a film and television review of the first video.

34. An electronic device, comprising: comprise: one or more processors; a display; a memory configured to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors and the display collectively implement the method of any one of claims 1-17.

35. A computer readable storage medium, characterized in that, The computer program, when executed on a computer, causes the computer to perform the method of any one of claims 1-17.

36. A computer program product, characterised in that, The computer program product comprises computer program code which, when executed on a computer, causes the computer to perform the method of any one of claims 1-17.

Citation Information

Patent Citations

  • Message pushing method and device, terminal, server and storage medium

    CN113965807A

  • Live broadcast question and answer and interface display method and computer storage medium

    CN114430490A

  • Video playing method and device

    CN116781971A

  • Question interaction method and device, computer equipment and storage medium

    CN117076642A

  • System and method for broadcast-synchronized interactive content interrelated to broadcast content

    CN1640129A