An interactive video playback system for interacting with characters in videos.

By receiving instruction information through the user interface module to generate dialogue data and interactive videos with virtual characters, the problem of video players being unable to interact is solved, thus improving the video viewing experience and interactivity.

CN116614665BActive Publication Date: 2026-01-30GUANGZHOU HONGRUI INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310639004.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-01
Publication Date
2026-01-30
Estimated Expiration
2043-06-01

AI Technical Summary

Technical Problem

Existing video players cannot enable interaction with characters in videos, resulting in a monotonous video viewing experience, especially in online teaching and live streaming scenarios where real-time interaction and efficient responses are not possible.

Method used

The system receives user commands through the user interface module, generates dialogue data, searches for matching video data and response text data from the knowledge base, constructs virtual characters, dynamically converts voice and actions, generates interactive videos of the virtual characters, and responds to user needs in real time.

Benefits of technology

It enables interaction with video characters, improves the video viewing experience, ensures the reliability and interactivity of video playback, and meets users' real-time needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116614665B_ABST
    Figure CN116614665B_ABST
Patent Text Reader

Abstract

This invention provides a video interactive playback system for interacting with characters in a video. It enables interactive video playback through a user interface module. During normal video playback, it generates matching dialogue data based on user-initiated commands and searches a knowledge base for matching video data and / or response text data. After processing the video data, it plays the real-time generated video to the user through the user interface module. Based on the characters in the video, it constructs virtual characters that interact with the user, and dynamically converts the virtual characters' voice and actions using response text data to generate interactive videos. These videos are then played through the user interface module, providing the user with a scenario interface for interacting with the virtual characters. The system receives and recognizes user-issued commands in real time, accurately grasps the user's video viewing needs, ensures reliable video playback, and can also create virtual characters matching the user for real-time video interaction, improving the user's video viewing experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent human-computer interaction, and particularly relates to a video interactive playing system for interacting with a character in a video. BACKGROUND

[0002] At present, a video player is a one-way video playing device, i.e., directly playing video images to a user. During the video playing process, the user cannot have an interactive conversation with a character existing in the video image, which makes the video watching experience single. In an online teaching or multimedia teaching scene, a student needs to ask questions at any time while watching a teaching video, but the teaching video is pre-recorded and cannot have a real-time teaching interaction with the student. In an online live broadcast or other commercial scene, in order to realize interaction with an audience, a real person is used to live broadcast, but for a large amount of interactive information from different audiences, the real person cannot fully and timely respond to all the interactive information, which also reduces the experience of live broadcast interaction and requires more human resources. During the video playing process, timely receiving and correctly recognizing the interactive request of the user and real-time and accurate returning of the interactive response to the user are the decisive factors for realizing video playing interaction. SUMMARY

[0003] In view of the defects in the prior art, the present application provides a video interactive playing system for interacting with a character in a video, which plays a video through a user interface module. During normal playing of the video, matched dialogue data is generated according to instruction information initiated by a user, and matched video data and / or response text data are searched from a knowledge base. After processing of the video data, the user interface module plays the real-time generated video to the user. According to the character in the video, a virtual character for interacting with the user is constructed, and the virtual character is dynamically converted in voice and action in combination with the response text data to generate a virtual character interactive video, which is played to the user through the user interface module to provide a scene interface for interacting with the virtual character. The instruction information initiated by the user is real-time received and recognized, the video watching demand of the user is correctly mastered, reliable video playing is ensured, a virtual character matched with the user can be created for real-time video interaction, and the video watching experience of the user is improved.

[0004] The present application provides a video interactive playing system for interacting with a character in a video, which comprises

[0005] a user interface module, which receives instruction information initiated by a user, and provides a scene interface for interacting with a virtual character and / or plays a video to the user;

[0006] The control module triggers the language recognition module and / or the video processing module to work according to the instruction information, and instructs the user interface module to form the scene interface for interacting with the virtual character according to the virtual character interactive video from the digital human module, and / or instructs the user interface module to play the video according to the playable video from the video processing module;

[0007] The language recognition module analyzes the instruction information to generate dialogue data matched with the instruction information.

[0008] The intelligent question and answer module finds matched video data and / or answer text data from the knowledge base according to the dialogue data.

[0009] The video processing module processes the video data to obtain a playable video, and sends the playable video to the control module.

[0010] The digital human module generates the virtual character interactive video according to the answer text data, and sends the virtual character interactive video to the control module.

[0011] Further, the user interface module before receiving the instruction information initiated by the user comprises:

[0012] The face image of the user currently interacting with the user interface module is captured, and the facial feature information of the user is extracted from the face image; whether the user is a legal user is determined according to the facial feature information; if the user is a legal user, the instruction information initiated by the user is received; if the user is not a legal user, the instruction information initiated by the user is not received.

[0013] Alternatively, the position information of the user currently interacting with the user interface module is detected, and whether the user is located within a predetermined activity range is determined according to the position information; if the user is located within the predetermined activity range, the instruction information initiated by the user is received; if the user is not located within the predetermined activity range, the instruction information initiated by the user is not received.

[0014] Further, the user interface module after receiving the instruction information initiated by the user further comprises:

[0015] The instruction information is sent to the control module to determine whether the instruction information is sound instruction information or text instruction information.

[0016] If the instruction information is a sound instruction information, then the sound instruction information is analyzed for useful sound signals and background sound signals to determine whether the sound instruction information is a valid sound instruction information; if it is a valid sound instruction information, then the content recognition of the valid sound instruction information is performed directly; if it is not a valid sound instruction information, then the user interface module is instructed to generate instruction information and resend the reminder message.

[0017] If the instruction information is text instruction information, then the text instruction information is subjected to text defect analysis processing to determine whether there are text errors in the text instruction information; if there are no text errors, then the text instruction information with text errors is directly subjected to content recognition; if there are text errors, then the user interface module is instructed to generate instruction information and resend the reminder message.

[0018] Furthermore, the control module triggers the language recognition module and / or video processing module to operate according to the instruction information, including:

[0019] The instruction information is subjected to content recognition to obtain the instruction code contained in the instruction information;

[0020] The instruction code is compared with a preset code directory, and the language recognition module and / or video processing module are triggered to work based on the comparison result.

[0021] Furthermore, the language recognition module analyzes the instruction information and generates dialogue data that matches the instruction information, including:

[0022] When the instruction information is a voice instruction information, the user's voice information components are extracted from the voice instruction information; the voice information components are subjected to speech recognition to generate text dialogue data that matches the instruction information;

[0023] When the instruction information is text instruction information, text dialogue data matching the instruction information is directly generated.

[0024] Furthermore, the intelligent question-answering module searches for matching video data and / or response text data from the knowledge base based on the dialogue data, including:

[0025] Extract all dialogue text words contained in the dialogue data, and generate a feature vector corresponding to the dialogue data based on all dialogue text words;

[0026] The feature vector is input into the knowledge base learning model to search for matching video data and / or response text data from the knowledge base;

[0027] The video data is sent to the video processing module and / or the response text data is sent to the digital human module.

[0028] The intelligent question-answering module can also call public AI analysis systems such as ChatGPT, Microsoft New Bing, Baidu Wenxin Yiyan, and Xunfei Xinghuo.

[0029] Furthermore, the video processing module processes the video data to obtain a playable video, and then sends the playable video to the control module, including:

[0030] The video data is segmented into frames to obtain several video image frames; the video image frames are then subjected to image content recognition and image restoration processing.

[0031] Then, based on the video playback parameters of the user interface module, the video data is converted into a video format to obtain a playable video; and the playable video is packaged, compressed, and sent to the control module.

[0032] The control module converts the playable video into a video stream playback signal and sends the video stream playback signal to the user interface module.

[0033] Furthermore, the video processing module performs content recognition on the video image frames and performs image restoration processing on the video image frames, including:

[0034] Step S1: Before performing image restoration processing, use the following formula (1) to determine whether the image content of the video image to be restored has been restored in other video images.

[0035]

[0036] In the above formula (1), X(b) represents the determination value of whether the content of the current b-th frame video image to be repaired has been repaired in other video images; ya(i,j) represents the pixel value of the pixel point in the i-th row and j-th column of the pixel matrix of the unrepaired a-th frame video image; y b (i, j) represents the pixel value of the pixel point in the i-th row and j-th column of the pixel matrix of the b-th frame of the video image to be repaired; | represents the absolute value; m represents the total number of pixels in any row of the pixel matrix of the video image; n represents the total number of pixels in any column of the pixel matrix of the video image.

[0037] If X(b) = 1, it means that the content of the b-th frame of the video image to be repaired has been repaired in other video images.

[0038] If X(b) = 0, it means that the content of the b-th frame of the video image to be repaired has not been repaired in other video images.

[0039] Step S2: If the image content has been repaired in other video images, then use the following formula (2) to locate the frame value of the image content in the other video images that have been repaired.

[0040]

[0041] In the above formula (2), A represents the image content of the b-th frame of the video image to be repaired in the array of frame values ​​of other video images that have been repaired; && represents the logical relationship AND; ∈ represents belonging to; This means that under the condition that a belongs to the interval [1, b-1], the following conditions will be met. All values ​​of 'a' are filtered out and arranged into an array in ascending order;

[0042] Step S3: Using the formula (3) below, verify whether the content of the repaired video image corresponding to the previous frame number is correct based on the frame value of the repaired video image in other video images.

[0043]

[0044] In the above formula (3), Z represents the determination value of whether the content of the b-th frame video image to be repaired is correctly repaired in the frame number of other video images that have been repaired.

[0045] If Z=0, then the content of the previously repaired frames needs to be repaired again.

[0046] If Z=1, then the content of the b-th frame video image to be repaired is directly controlled to be repaired into the repaired content corresponding to any element in array A.

[0047] Furthermore, the digital human module generates an interactive video of the virtual character based on the response text data, and sends the interactive video of the virtual character to the control module, including:

[0048] Based on the user's video interaction history logs, construct a corresponding virtual character and set the characteristic elements of the virtual character's voice and action behaviors;

[0049] Based on the response text data, virtual character voice information is generated; then, based on the virtual character and the virtual character voice information, the virtual character interactive video is generated; and the virtual character interactive video is packaged, compressed, and sent to the control module.

[0050] The control module converts the virtual character's interactive video into a video stream playback signal and sends the video stream playback signal to the user interface module.

[0051] Compared to existing technologies, this interactive video playback system for interacting with characters in videos uses a user interface module for interactive video playback. During normal video playback, it generates matching dialogue data based on user-initiated commands and searches for matching video data and / or response text data from a knowledge base. After processing the video data, it plays the real-time generated video to the user through the user interface module. Based on the characters in the video, it constructs virtual characters that interact with the user and dynamically converts the virtual characters' voice and actions using response text data to generate interactive videos. These videos are then played through the user interface module, providing users with a scenario interface for interacting with virtual characters. The system receives and recognizes user-issued commands in real time, accurately grasps the user's video viewing needs, ensures reliable video playback, and can also create virtual characters that match the user for real-time video interaction, improving the user's video viewing experience.

[0052] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings.

[0053] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0055] Figure 1 This is a schematic diagram of the structure of the video interactive playback system provided by the present invention for interacting with characters in a video. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] See Figure 1 This is a schematic diagram of a video interactive playback system for interacting with characters in a video, provided in an embodiment of the present invention. The video interactive playback system for interacting with characters in a video includes:

[0058] The user interface module receives user-initiated commands and provides users with a scene interface for interacting with virtual characters and / or plays videos.

[0059] The control module, based on the instruction information, triggers the language recognition module and / or the video processing module to work; and, based on the virtual character interaction video from the digital human module, instructs the user interface module to form the scene interface for interacting with the virtual character, and / or, based on the playable video from the video processing module, instructs the user interface module to play the video.

[0060] The language recognition module analyzes the instruction information and generates dialogue data that matches the instruction information;

[0061] The intelligent question-answering module searches for matching video data and / or response text data from the knowledge base based on the dialogue data.

[0062] The video processing module processes the video data to obtain a playable video, and then sends the playable video to the control module.

[0063] The digital human module generates an interactive video of the virtual character based on the response text data, and sends the interactive video of the virtual character to the control module.

[0064] The beneficial effects of the above technical solution are as follows: This video interactive playback system for interacting with characters in a video interacts with the user interface module. During normal video playback, it generates matching dialogue data based on user-initiated commands and searches for matching video data and / or response text data from a knowledge base. After processing the video data, it plays the real-time generated video to the user through the user interface module. Based on the characters in the video, it constructs virtual characters that interact with the user and, combined with response text data, dynamically converts the virtual characters' voice and actions to generate virtual character interactive videos. These videos are then played through the user interface module, providing users with a scenario interface for interacting with virtual characters. The system receives and recognizes user-issued commands in real time, accurately grasps the user's video viewing needs, ensures reliable video playback, and can also create virtual characters that match the user for real-time video interaction, improving the user's video viewing experience.

[0065] Preferably, before receiving user-initiated command information, the user interface module includes:

[0066] Capture a facial image of the user currently interacting with the user interface module, and extract the user's facial feature information from the facial image; determine whether the user is a legitimate user based on the facial feature information; if the user is a legitimate user, receive the instruction information initiated by the user; if the user is not a legitimate user, do not receive the instruction information initiated by the user.

[0067] Alternatively, the location information of the user currently interacting with the user interface module can be detected, and based on the location information, it can be determined whether the user is within a predetermined activity range; if the user is within the predetermined activity range, the instruction information initiated by the user can be received; if the user is not within the predetermined activity range, the instruction information initiated by the user can not be received.

[0068] The beneficial effects of the above technical solution are as follows: The user interface module may include, but is not limited to, a display with sound playback and sound reception functions, and a camera mounted on the display. The display provides a scene interface for users to interact with virtual characters and / or plays videos; the display may be, but is not limited to, a touch screen with speakers and a microphone. The camera is used to capture images of the user currently interacting with the user interface module. The camera may be, but is not limited to, a binocular camera. The binocular camera captures images of the user's face, obtaining a binocular image of the user's face. Based on the binocular parallax of the user's facial binocular image, a three-dimensional facial image is generated. Then, facial feature contour recognition processing is performed on the three-dimensional facial image to obtain the user's facial feature contour information. The facial contour features are compared with a preset facial feature database. If the facial contour features exist in the database, the user is considered a legitimate user, and the system accepts user input via microphone and / or touchscreen. If the facial contour features do not exist in the database, the user is considered illegitimate, and the microphone and touchscreen are instructed to stop working, thus refusing to accept user input. This ensures the security of the video interactive playback system. Furthermore, a binocular camera can be used to capture images of the user in their current environment. Parallax analysis is performed on these images to determine the relative distance and orientation between the user and the user interface module. Based on this relative distance and orientation, it is determined whether the user is currently within a predetermined activity range, which is a predetermined area near the user interface module. When the user is within the designated activity range, it indicates that the user interface module can accurately receive the instructions issued by the user. At this time, it receives the instructions input by the user through the microphone and / or touch screen. When the user is not within the designated activity range, it indicates that the user interface module cannot accurately receive the instructions issued by the user. At this time, it instructs the microphone and touch screen to stop working, thereby refusing to receive the instructions input by the user. This can ensure the reliability of the video interactive playback system.

[0069] Preferably, after receiving the user-initiated instruction information, the user interface module further includes:

[0070] The instruction information is sent to the control module to determine whether it is an audio instruction or a text instruction.

[0071] If the instruction information is an audio instruction, the useful audio signal and background audio signal are analyzed to determine whether the audio instruction information is a valid audio instruction. If it is a valid audio instruction, the content of the valid audio instruction is directly recognized. If it is not a valid audio instruction, the user interface module is instructed to generate instruction information and resend the reminder message.

[0072] If the instruction is a text instruction, then the text instruction is subjected to text defect analysis to determine whether there are any text errors. If there are no text errors, then the text instruction is directly subjected to content recognition. If there are text errors, then the user interface module is instructed to generate instruction information and resend the reminder message.

[0073] The beneficial effects of the above technical solution are as follows: Users can input command information to the user interface module via voice or text. When the control module receives the command information from the user, it first identifies whether the command information is a voice command or a text command, thus enabling precise differentiation. When the command information is a voice command, the system extracts the user's voice signal and the background sound signal generated by the user's environment based on the user's voiceprint characteristics. Then, based on the useful voice signal and the background sound signal, the signal-to-noise ratio (SNR) of the voice command information is determined. If the SNR is greater than or equal to a preset SNR threshold, the voice command information is considered valid. In this case, content recognition is directly performed on the voice command information to obtain the contained command content, thereby improving the accuracy of command content recognition. When the command information is a text command, the system performs text analysis to determine if there are any typos or other text errors. If no text errors are found, content recognition is directly performed on the text command information to obtain the contained command content, thereby improving the accuracy of command content recognition.

[0074] Preferably, the control module triggers the language recognition module and / or video processing module to operate according to the instruction information, including:

[0075] The instruction information is subjected to content recognition to obtain the instruction code contained in the instruction information;

[0076] The instruction code is compared with the preset code directory, and the language recognition module and / or video processing module are triggered based on the comparison result.

[0077] The beneficial effects of the above technical solution are as follows: After content recognition of the instruction information, the semantic content of the instruction information can be obtained, which also contains the corresponding instruction code. This instruction code is then compared with a preset code directory, which includes several predetermined instruction codes corresponding to the language recognition module and the video processing module, respectively. If the instruction code matches the predetermined instruction codes corresponding to the language recognition module and the video processing module, the language recognition module and / or the video processing module are triggered to operate. This ensures that the language recognition module and the video processing module only operate under appropriate conditions, avoiding increasing the computational workload of the language recognition module and the video processing module.

[0078] Preferably, the language recognition module analyzes the instruction information and generates dialogue data that matches the instruction information, including:

[0079] When the instruction information is a voice instruction information, the user's voice information component is extracted from the voice instruction information; the voice information component is then subjected to speech recognition to generate text dialogue data that matches the instruction information;

[0080] When the instruction information is a text instruction information, text dialogue data that matches the instruction information is directly generated.

[0081] The beneficial effects of the above technical solution are as follows: For different situations where the instruction information is voice instruction information and text instruction information, the language recognition module needs to perform different working modes of speech recognition and text recognition. That is, the language recognition module has both speech recognition and text recognition functions. This can ensure that the user's needs can be accurately identified when the user inputs different forms of instruction information, and achieve multi-scenario matching of user instruction recognition.

[0082] Preferably, the intelligent question-answering module searches for matching video data and / or response text data from a knowledge base based on the dialogue data, including:

[0083] Extract all the dialogue text words contained in the dialogue data, and generate the feature vector corresponding to the dialogue data based on all the dialogue text words.

[0084] The feature vector is input into the knowledge base learning model to search for matching video data and / or response text data from the knowledge base;

[0085] Send the video data to the video processing module and / or send the response text data to the digital human module.

[0086] The beneficial effects of the above technical solution are as follows: When corresponding dialogue data is obtained from the user's input command information, the dialogue text vocabulary contained in the dialogue data is processed by models such as ChatGPT, Microsoft New Bing, Baidu Wenxin Yiyan, and Xunfei Xinghuo to generate a feature vector corresponding to the dialogue data. This feature vector can be, but is not limited to, a vector obtained by mathematically transforming the dialogue text vocabulary; this is a commonly used vocabulary transformation method in neural network models, which will not be described in detail here. The feature vector is then input into a knowledge base learning model to search for matching video data and / or response text data in the knowledge base; this knowledge base includes different types of knowledge data, which can be in text form and / or video form. The video data is then sent to the video processing module and / or the response text data is sent to the digital human module, enabling the video processing module and the digital human module to perform corresponding data processing.

[0087] Preferably, the video processing module processes the video data to obtain a playable video, and then sends the playable video to the control module, including:

[0088] The video data is segmented into frames to obtain several video image frames; the content of the video image frames is identified, and the video image frames are then repaired.

[0089] Then, based on the video playback parameters of the user interface module, the video data is converted into a video format to obtain a playable video; and the playable video is packaged, compressed, and sent to the control module.

[0090] The control module converts the playable video into a video stream playback signal and sends the video stream playback signal to the user interface module.

[0091] The beneficial effects of the above technical solution are as follows: by performing frame-by-frame processing and image restoration on the video data through the video processing module, the image quality of the video data can be improved. Furthermore, by converting the video data into a playable video based on the video playback parameters of the user interface module, the user interface module can ensure that the video playback is normal and smooth, avoid video playback stuttering, and improve the user's video viewing experience.

[0092] Preferably, the video processing module performs content recognition on the video image frame and image restoration processing on the video image frame, including:

[0093] Step S1: Before performing image restoration processing, use the following formula (1) to determine whether the image content of the video image to be restored has been restored in other video images.

[0094]

[0095] In the above formula (1), X(b) represents the judgment value of whether the content of the b-th frame of the video image to be repaired is among the content of other video images that have been repaired; y a (i, j) represents the pixel value of the pixel point in the i-th row and j-th column of the pixel matrix of the unrepaired a-th frame of the video image; y b (i, j) represents the pixel value of the pixel in the i-th row and j-th column of the pixel matrix of the b-th frame of the video image to be repaired; | represents the absolute value; m represents the total number of pixels in any row of the pixel matrix of the video image; n represents the total number of pixels in any column of the pixel matrix of the video image.

[0096] If X(b) = 1, it means that the content of the b-th frame of the video image to be repaired has been repaired in other video images.

[0097] If X(b) = 0, it means that the content of the b-th frame of the video image to be repaired has not been repaired in other video images.

[0098] Step S2: If the content of the image has been repaired in other video images, then use the following formula (2) to locate the frame value of the content of the image in the other video images that have been repaired.

[0099]

[0100] In the above formula (2), A represents the image content of the b-th frame of the video image to be repaired in the array of frame values ​​of other video images that have been repaired; && represents the logical relationship AND; ∈ represents belonging to; This means that under the condition that a belongs to the interval [1, b-1], the following conditions will be met. All values ​​of 'a' are filtered out and arranged into an array in ascending order;

[0101] Step S3: Using formula (3) below, verify whether the content of the repaired frame corresponding to the previous frame number is correct based on the frame value of the repaired frame in other video images.

[0102]

[0103] In the above formula (3), Z represents the determination value of whether the content of the b-th frame video image to be repaired is correctly repaired in the frame number corresponding to the repaired frame of other video images.

[0104] If Z=0, then the content of the previously repaired frames needs to be repaired again.

[0105] If Z=1, then the content of the b-th frame of the video image to be repaired is directly controlled to be repaired and processed into the repaired content corresponding to any element in array A.

[0106] The beneficial effects of the above technical solution are as follows: Using the above formula (1), the screen content of the video image to be repaired is determined to see if the screen content has been repaired in other video images, thus quickly checking whether it has been repaired, which reflects the system's ability to quickly check; then using the above formula (2), the frame value of the screen content that has been repaired in other video images is located, so that the result of its previous repair can be known and the accuracy of its repair can be verified, ensuring the accuracy of the system; finally using the above formula (3), the screen content that has been repaired in other video images is verified to see if the repaired screen content corresponding to the previous frame number is correct, and after repeated verification, the overall reliability of the system is ensured.

[0107] Preferably, the digital human module generates an interactive video of the virtual character based on the response text data, and sends the interactive video of the virtual character to the control module, including:

[0108] Based on the user's video interaction history logs, construct a corresponding virtual character and set the characteristic elements of the virtual character's voice and action behaviors;

[0109] Based on the response text data, generate virtual character voice information; then, based on the virtual character and the virtual character voice information, generate virtual character interaction video; and package and compress the virtual character interaction video and send it to the control module.

[0110] The control module converts the virtual character's interactive video into a video stream playback signal and sends the video stream playback signal to the user interface module.

[0111] The beneficial effects of the above technical solution are as follows: The digital human module can be, but is not limited to, a virtual image construction module, which has a virtual animation generation function, thereby generating video animations containing virtual characters. It extracts the virtual character images that the user habitually uses during historical video interactions from the user's video interaction history logs, constructs corresponding virtual characters, and then sets the characteristic elements of the virtual character's voice and action behaviors. The voice behavior characteristic elements can include, but are not limited to, the tone, volume, and speed of the virtual character's voice; the action behavior characteristic elements can include, but are not limited to, the type and amplitude of the virtual character's body movements during speech. This enhances the realism of the user's interaction with the virtual character during video interaction. Furthermore, the virtual character interaction video can be integrated and overlaid with the playable video, allowing the user to interact with the virtual character in real time while watching the video.

[0112] As described in the above embodiments, this video interactive playback system for interacting with characters in a video interacts with the user interface module. During normal video playback, it generates matching dialogue data based on user-initiated commands and searches for matching video data and / or response text data from a knowledge base. After processing the video data, it plays the real-time generated video to the user through the user interface module. Based on the characters in the video, it constructs virtual characters that interact with the user, and combines this with response text data to dynamically convert the virtual characters' voice and actions, generating interactive videos of these virtual characters. These videos are then played through the user interface module, providing the user with a scenario interface for interacting with the virtual characters. The system receives and recognizes user-issued commands in real time, accurately grasps the user's video viewing needs, ensures reliable video playback, and can also create virtual characters matching the user for real-time video interaction, improving the user's video viewing experience.

[0113] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A video interactive playing system for interacting with a character in a video, comprising: a user interface module, receiving user-initiated instruction information, and providing a user with a scenario interface for interacting with a virtual character and / or playing a video; a control module, triggering a language recognition module and / or a video processing module to work according to the instruction information; and instructing the user interface module to form the scenario interface for interacting with the virtual character according to a virtual character interactive video from a digital human module, and / or instructing the user interface module to play a video according to a playable video from the video processing module; the language recognition module, analyzing the instruction information to generate dialogue data matching the instruction information; an intelligent question and answer module, searching for matching video data and / or answer text data from a knowledge base according to the dialogue data; the video processing module, processing the video data to obtain a playable video, and sending the playable video to the control module; frame processing the video data to obtain a plurality of video image frames; picture content recognition and picture repair processing of the video image frames, comprising: step S1, before picture repair processing, judging whether the picture content of the current video image to be repaired has been repaired in other video images; step S2, if the picture content has been repaired in other video images, locating the frame number of the picture content that has been repaired in other video images; step S3, according to the frame number of the picture content that has been repaired in other video images, verifying whether the repaired picture content corresponding to the previous frame number is correct; if not, the repaired picture content corresponding to the previous frame number needs to be repaired again; if correct, the picture content of the video image to be repaired is directly repaired to the corresponding repaired picture content; after repair processing, video format conversion of the video data according to video playing parameters of the user interface module to obtain a playable video; and packaging and compressing the playable video and sending it to the control module; the control module converts the playable video into a video stream playing signal and sends the video stream playing signal to the user interface module; the digital human module generates the virtual character interactive video according to the answer text data and sends the virtual character interactive video to the control module. 2.The video interactive playing system for interacting with a character in a video according to claim 1, wherein: before receiving the user-initiated instruction information, the user interface module comprises: capturing a face image of a user currently interacting with the user interface module, extracting face feature information of the user from the face image, judging whether the user is a legal user according to the face feature information, if the user is a legal user, receiving the user-initiated instruction information, if the user is not a legal user, not receiving the user-initiated instruction information. Alternatively, the position information of a user currently interacting with the user interface module is detected, and whether the user is located within a predetermined activity range is determined according to the position information; if the user is located within the predetermined activity range, the instruction information initiated by the user is received; and if the user is not located within the predetermined activity range, the instruction information initiated by the user is not received.

3. The video interactive playing system for interacting with a character role in a video according to claim 1, characterized in that: after the user interface module receives the instruction information initiated by the user, the user interface module further comprises: sending the instruction information to the control module to determine whether the instruction information belongs to sound instruction information or text instruction information; if the instruction information belongs to sound instruction information, performing useful sound signal and background sound signal analysis on the sound instruction information to determine whether the sound instruction information belongs to valid sound instruction information; if the sound instruction information belongs to valid sound instruction information, directly performing content recognition on the valid sound instruction information; if the sound instruction information does not belong to valid sound instruction information, instructing the user interface module to generate an instruction information re-sending reminder message; if the instruction information belongs to text instruction information, performing text defect analysis processing on the text instruction information to determine whether the text instruction information has a text error; if the text instruction information does not have a text error, directly performing content recognition on the text instruction information; if the text instruction information has a text error, instructing the user interface module to generate an instruction information re-sending reminder message.

4. The video interactive playing system for interacting with a character role in a video according to claim 1, characterized in that: the control module triggers the language recognition module and / or the video processing module to work according to the instruction information, comprising: performing content recognition on the instruction information to obtain instruction codes contained in the instruction information; comparing the instruction codes with a preset code directory, and triggering the language recognition module and / or the video processing module to work according to the comparison result.

5. The video interactive playing system for interacting with a character role in a video according to claim 1, characterized in that: the language recognition module analyzes the instruction information to generate dialogue data matched with the instruction information, comprising: when the instruction information belongs to sound instruction information, extracting speech information components of the user from the sound instruction information; performing speech recognition on the speech information components to generate text dialogue data matched with the instruction information; when the instruction information belongs to text instruction information, generating text dialogue data matched with the instruction information.

6. The video interactive playing system for interacting with a character role in a video according to claim 1, characterized in that: the intelligent question and answer module finds matched video data and / or answer text data from a knowledge base according to the dialogue data, and the intelligent question and answer module can call APIs of ChatGPT, Microsoft New Bing, Baidu Ernie Bot, and Xunfei Starfire Platform. sending the video data to the video processing module and / or sending the answer text data to the digital human module. 7.The video interactive playing system for interacting with a character in a video according to claim 1, wherein: the intelligent question and answer module finds matching video data and / or answer text data from a knowledge base according to the dialogue data, including: extracting all dialogue text vocabularies contained in the dialogue data, and generating a feature vector corresponding to the dialogue data according to all dialogue text vocabularies; inputting the feature vector into a knowledge base learning model to find matching video data and / or answer text data from the knowledge base; sending the video data to the video processing module and / or sending the answer text data to the digital human module. 8.The video interactive playing system for interacting with a character in a video according to claim 1, wherein: the step S1, before performing the picture inpainting processing, judges whether the picture content of the current video image to be processed is in other video images which are processed, by using the following formula (1), ; In the above formula (1), X(b) represents a determination value of whether the picture content of the bth frame video image currently to be repaired is repaired in other video images; y a (i,j) represents the pixel value of the i-th row and j-th column pixel point in the pixel matrix of the a-th frame video image which is not repaired; y b (i,j) represents the pixel value of the i-th row and j-th column pixel point in the pixel matrix of the b-th frame video image currently to be repaired; || represents the absolute value; m represents the total number of any one row of pixel points in the pixel matrix of the video image; n represents the total number of any one column of pixel points in the pixel matrix of the video image; if X(b) = 1, it indicates that the picture content of the b-th frame of video image to be processed is in other video images which are processed; if X(b) = 0, it indicates that the picture content of the b-th frame of video image to be processed is not in other video images which are processed; the step S2, if the picture content is in other video images which are processed, then using the following formula (2), locates the frame value of the picture content in other video images which are processed, ; In the above formula (2), A represents the image content of the bth frame video image to be repaired in the frame number array in which other video images are repaired; & represents a logical relationship; and ∈ represents belonging to; represents that all a values satisfying are screened out and arranged into an array in ascending order. the step S3, using the following formula (3), according to the frame value of the picture content in other video images which are processed, verifies whether the picture content corresponding to the previous frame is correct, ; in the above formula (3), Z represents the judgment value of whether the picture content corresponding to the frame value of the picture content in other video images which are processed is processed correctly; if Z = 0, it is necessary to reprocess the picture content corresponding to the previous frame which is processed; if Z = 1, it directly controls the picture content of the b-th frame of video image to be processed to be processed into the picture content corresponding to any one element in the array A. 9.The video interactive playing system for interacting with a character in a video according to claim 1, wherein: the digital human module generates the virtual character interactive video according to the answer text data, and sends the virtual character interactive video to the control module, including: according to the video interactive history log of the user, constructing a corresponding virtual character, and setting the feature elements of the virtual character to perform voice behavior and action behavior; According to the response text data, virtual character voice information is generated; according to the virtual character and the virtual character voice information, the virtual character interactive video is generated; and the virtual character interactive video is packaged, compressed and sent to the control module; The control module converts the virtual character interactive video into a video stream playing signal, and sends the video stream playing signal to the user interface module.

Citation Information

Patent Citations

  • Video image quality enhancement method and related equipment thereof

    CN110996174A

  • Interactive video playing method and system

    CN111970568A

  • Role interaction method and device, electronic equipment and storage medium

    CN116072095A