Live video generation method, live picture display method, device and system
By identifying the action, expression and scene information in the live video stream, using the pre-trained model to generate a matching comic-style video stream, and allowing the audience to select update elements, the problem of single visual effects of the live video stream is solved, and rich visual experience and high interactivity are achieved.
Patent Information
- Application Number
- CN202510475552.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-11
AI Technical Summary
The existing live video streams are relatively single in terms of visual effects and lack innovation. The anchor needs to actively operate and add special effects elements and cannot dynamically adapt, which lacks flexibility.
By identifying the character's actions, expressions or scene information in the original video stream, the pre-trained comic style conversion model generates a matching target video stream, and the comic elements can be updated according to the audience's selection to achieve automated dynamic generation of comic styles.
It enriches the visual effects of the live broadcast screen, improves the fun and interactiveness of the live broadcast content, reduces manual intervention, and the generated comic style is more natural and vivid.
Smart Images

Figure CN120302105A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of network live broadcast technology, and particularly to a live video generation method, a live video display method, a device and a system. Background Art
[0002] With the rapid development of the live broadcast industry, watching live broadcasts has become an important way for users to relax and entertain. Currently, the live video stream is usually the original video stream directly obtained by the video stream acquisition device (such as a camera device) at the anchor end, and the visual effect presentation is relatively single and lacks innovation. In some live broadcast platforms, although the anchor can actively select some special effect elements to add to the live video during the live broadcast, this requires the active operation of the anchor, and the added special effect elements are fixed and cannot be dynamically adapted based on the content of the live video, lacking flexibility.
[0003] In view of this, some embodiments of this specification provide a live video generation method, a live video display method, a device and a system, aiming to automatically and dynamically generate a target video stream with a comic style that matches the content of the original video stream, enrich the visual effect of the live video, and enhance the interest of the live content. Summary of the Invention
[0004] One or more embodiments of this specification provide a live video generation method, the method includes: obtaining an original video stream; identifying first information in the original video stream, the first information including at least one of human action information, human expression information, or scene information in the original video stream; converting the original video stream based on the first information to generate a target video stream with a target comic style that matches the original video stream.
[0005] According to the method provided by one or more embodiments of this specification, it further includes: obtaining selection information of candidate comic elements sent by the viewer end, the candidate comic elements including at least one of candidate comic styles, candidate comic special effects, or candidate comic characters; determining updated target comic elements based on the selection information of the candidate comic elements, and generating an updated target video stream based on the updated target comic elements.
[0006] According to the method provided by one or more embodiments of this specification, when the selection information of the candidate comic elements includes a candidate comic style, determining updated target comic elements based on the selection information of the candidate comic elements, and generating an updated target video stream based on the updated target comic elements, includes: determining an updated target comic style based on the selection information of the candidate comic style; converting the original video stream based on the updated target comic style to generate an updated target video stream, and the updated target video stream has the updated target comic style.
[0007] The method provided according to one or more embodiments of this specification, based on the selection information of candidate comic elements, determines updated target comic elements, including: based on the statistical quantity of the selection information of candidate comic elements, determining the candidate comic element with the largest selection quantity as the updated target comic element; or based on the statistical quantity of the selection information of candidate comic elements, determining the weight information of each candidate comic element, and generating a comprehensive comic element as the updated target comic element, where the comprehensive comic element integrates the characteristics of each candidate comic element.
[0008] The method provided according to one or more embodiments of this specification, when the selection information of candidate comic elements includes candidate comic special effects or candidate comic characters, based on the selection information of candidate comic elements, determines updated target comic elements, and generates an updated target video stream based on the updated target comic elements, including: based on the selection information of candidate comic special effects or candidate comic characters, determining target comic special effects or target comic characters; adjusting the target video stream with a target comic style based on the target comic special effects or target comic characters to generate an updated target video stream, and the updated target video stream includes the target comic special effects or target comic characters.
[0009] The method provided according to one or more embodiments of this specification, adjusting the target video stream with a target comic style based on the target comic special effects or target comic characters to generate an updated target video stream, including: adding target comic special effects to the target video stream with a target comic style; or replacing the comic special effects in the target video stream with a target comic style with the target comic special effects; or replacing the characters in the target video stream with a target comic style with the target comic characters.
[0010] The method provided according to one or more embodiments of this specification further includes: determining candidate comic elements based on preset information or historical selection information of the viewer terminal for the viewer terminal to select based on the interactive interface.
[0011] The method provided according to one or more embodiments of this specification, the human motion information includes human body key point information and / or human motion categories; identifying the first information in the original video stream includes: identifying the human body key point information in the original video stream; determining the human motion categories in the original video stream based on the human body key point information.
[0012] The method provided according to one or more embodiments of this specification, the human expression information includes at least one of human expression categories, facial key point information, and facial feature vectors; identifying the first information in the original video stream includes: identifying the facial key point information and / or facial feature vectors in the original video stream; determining the human expression categories in the original video stream based on the facial key point information and / or facial feature vectors.
[0013] According to the method provided by one or more embodiments of this specification, the scene information includes at least one of person information, object information, and background information; the person information includes person position and person category, the object information includes object position and object category, and the background information includes background area and background category; wherein, the person position includes human key point information; the object position includes the bounding box information of the object; the background area is determined based on the mask image of the background.
[0014] According to the method provided by one or more embodiments of this specification, the original video stream is sent from the anchor end; the person action information includes the action information of the anchor; the person expression information includes the expression information of the anchor; the scene information includes the scene information of the live broadcast room where the anchor is located.
[0015] According to the method provided by one or more embodiments of this specification, it further includes: adding a comic special effect matching the first information to the target video stream based on the first information.
[0016] According to the method provided by one or more embodiments of this specification, the conversion of the original video stream based on the first information is achieved through a pre-trained comic style conversion model; the pre-trained comic style conversion model generates a target video stream with a target comic style matching the original video stream based on the first information and each original video frame in the original video stream, or based on the fusion feature map or fusion feature vector generated from the first information and each original video frame of the original video stream.
[0017] According to the method provided by one or more embodiments of this specification, the comic style conversion model is trained based on a sample training set, and the sample training set includes video frame samples, first sample information extracted from the video frame samples, and comic style images matching the first sample information, and the first sample information includes at least one of action sample information, expression sample information, and scene sample information.
[0018] One or more embodiments of this specification also provide a live broadcast screen display method, and the method includes: receiving a target video stream with a target comic style sent by the server side, the target video stream is generated by the server side based on the conversion of the original video stream in the original video stream sent from the anchor end, and the target video stream has a target comic style matching the original video stream; displaying the target video stream with the target comic style on the live broadcast screen; wherein, the first information includes at least one of person action information, person expression information, or scene information in the original video stream.
[0019] The method provided according to one or more embodiments of this specification further includes: displaying candidate comic elements, where the candidate comic elements include at least one of a candidate comic style, candidate comic special effects, or candidate comic characters; obtaining selection information of the candidate comic elements input based on an interaction interface and sending it to the server side; receiving and displaying the updated target video stream sent by the server side, where the updated target video stream is generated by the server side based on the selection information of the candidate comic elements.
[0020] For the method provided according to one or more embodiments of this specification, when the selection information of the candidate comic elements includes a candidate comic style, the updated target video stream has an updated target comic style determined based on the selection information of the candidate comic style; when the selection information of the candidate comic elements includes candidate comic special effects or candidate comic characters, the updated target video stream includes target comic special effects or target comic characters determined based on the selection information of the candidate comic special effects or candidate comic characters.
[0021] For the method provided according to one or more embodiments of this specification, the updated target video stream is generated by the server side based on the updated target comic elements, where the updated target comic elements are the candidate comic elements with the largest number of selections or composite comic elements, and the composite comic elements are generated based on the weight information of each candidate comic element.
[0022] One or more embodiments of this specification also provide a live broadcast screen display method, and the method includes: obtaining and sending the original video stream of the live broadcast screen to the server side; receiving the target video stream with the target comic style sent by the server side, where the target video stream is generated by the server side based on the first information in the original video stream to convert the original video stream and has a target comic style matching the original video stream; displaying the target video stream in the live broadcast screen; where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream.
[0023] One or more embodiments of this specification also provide a live video generation device, and the device includes: an obtaining module for obtaining the original video stream; an identifying module for identifying the first information in the original video stream, where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream; a conversion module for converting the original video stream based on the first information to generate a target video stream with a target comic style matching the original video stream.
[0024] One or more embodiments of this specification also provide a live video display device, which includes: a receiving module, configured to receive a target video stream in a target comic style sent by a server end, where the target video stream is generated by the server end based on first information in the original video stream sent by a live broadcast end to convert the original video stream, and the target video stream has a target comic style matching the original video stream; a display module, configured to display the target video stream in the target comic style in the live video; where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream.
[0025] One or more embodiments of this specification also provide a live video display device, which includes: an acquisition module, configured to acquire and send the original video stream of the live video to the server end; a receiving module, configured to receive a target video stream in a target comic style sent by the server end, where the target video stream is generated by the server end based on first information in the original video stream to convert the original video stream, and has a target comic style matching the original video stream; a display module, configured to display the target video stream in the live video; where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream.
[0026] One or more embodiments of this specification also provide a live video system, which includes a server end, a live broadcast end, and an audience end; the server end is configured to acquire the original video stream sent by the live broadcast end, identify the first information in the original video stream, convert the original video stream based on the first information, generate a target video stream with a target comic style matching the original video stream, and send it to the live broadcast end and the audience end; the live broadcast end is configured to acquire and send the original video stream to the server end, and acquire and display the target video stream in the target comic style sent by the server end; the audience end is configured to acquire and display the target video stream in the target comic style sent by the server end; where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream.
[0027] One or more embodiments of this specification also provide a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it can implement the live video generation method and the live video display method described in some embodiments of this specification.
[0028] One or more embodiments of this specification also provide a computer-readable storage medium, which stores computer instructions. When the computer instructions are executed by a processor, they can implement the live video generation method and the live video display method described in some embodiments of this specification.
[0029] One or more embodiments of this specification also provide a computer program product, including a computer program, which can implement the live video generation method and the live video display method described in some embodiments of this specification when at least a part of the computer program is executed by a processor. Description of the Drawings
[0030] This specification will be further described by way of exemplary embodiments, which will be described in detail through the drawings. The same numbers in the drawings represent the same structures or steps.
[0031] Figure 1 It is a schematic diagram of a live operation environment shown in some embodiments of this specification.
[0032] Figure 2 It is an exemplary flowchart of a live video generation method shown in some embodiments of this specification.
[0033] Figure 3 It is a schematic diagram of a video frame of an original video stream shown in some embodiments of this specification.
[0034] Figure 4 It is a schematic diagram of a video frame of a target video stream shown in some embodiments of this specification.
[0035] Figure 5 It is an exemplary flowchart of another live video generation method shown in some embodiments of this specification.
[0036] Figure 6 It is a schematic diagram of a video frame of an updated target video stream shown in some embodiments of this specification.
[0037] Figure 7 It is an exemplary flowchart of a live video display method shown in some embodiments of this specification.
[0038] Figure 8 It is an exemplary flowchart of another live video display method shown in some embodiments of this specification.
[0039] Figure 9 It is an exemplary flowchart of a live video display method shown in some embodiments of this specification.
[0040] Figure 10 It is an exemplary block diagram of a live video generation device shown in some embodiments of this specification.
[0041] Figure 11 It is an exemplary block diagram of a live video display device shown in some embodiments of this specification.
[0042] Figure 12 It is an exemplary block diagram of a live video display device shown in some embodiments of this specification.
[0043] Figure 13 It is an exemplary block diagram of a live broadcast system shown in some embodiments of this specification. Detailed implementation manners
[0044] To more clearly illustrate the technical solutions of the embodiments of this specification, the embodiments will be introduced in detail below with reference to the accompanying drawings. Obviously, the content described below is some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, the technical solutions or means disclosed in this specification can also be applied to other scenarios according to these technical contents.
[0045] It should be understood that the "system", "device", "unit" and / or "module" used in this specification is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the said words can be replaced by other expressions.
[0046] Unless otherwise specified, the technical terms describing components, elements, etc. in this specification do not specifically refer to the singular, but may also include the plural. Generally speaking, terms such as "including" and "comprising" only indicate the inclusion of the clearly identified steps, elements or components, and these steps, elements and components do not constitute an exclusive list. For example, the described method or device may also include other steps or components.
[0047] Flowcharts are used in this specification to illustrate the operation steps performed by the devices or systems of the relevant embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of step execution. Those of ordinary skill in the art can adjust the order of execution of these steps according to the knowledge information conveyed by the embodiments of this specification. The said adjustments include, but are not limited to, the swapping of the sequence, the merging of multiple steps, and the splitting of a certain step.
[0048] Live broadcast (or called webcast) is a technical form of real-time transmission of audio and video content over the Internet. The host side can collect the live video of the host in real time through a camera device, and upload the video data in the form of video streams, video files, data packets for chunked transmission, etc. to the server side, and the server side sends it to the viewer side. Users on the viewer side can watch the live content online and can interact with the host in real time. For example, users on the viewer side can interact with the host through bullet screens, rewards, voice calls, etc.
[0049] Figure 1It is a schematic diagram of a live broadcast operating environment shown according to some embodiments of this specification. As Figure 1 shown, the live broadcast operating environment 100 may include: a server side 110, a host side 120, an audience side 130, and a network 140. The server side 110, the host side 120, and the audience side 130 may perform data transmission through the network 140. During the live broadcast, the host side 120 may send the video data of the live broadcast to the server side 110 through the network 140, and the server side 110 may further process the video data of the live broadcast and send it to the audience side 130 through the network 140 for the users of the audience side 130 to watch.
[0050] The server end 110 may be a computer device with high computing performance, used for storing, forwarding, configuring, analyzing and processing video data, etc. For example, the server end 110 may forward the original video data (e.g., original video stream) sent by the anchor end 120, or the server end 110 may further analyze and process the original video data and send it to the audience end 130; for another example, the server end 110 may process the interactive data (e.g., likes, bullet screens, or interactive data for changing the style, special effects, character images, etc. of the live screen) sent by the audience end 130, and send the processed video data (e.g., video stream) to the audience end 130 and the anchor end 120. In some embodiments, the server end 110 may transmit data based on a streaming media transmission protocol, and the streaming media transmission protocol may include: HTTP Live Streaming (HLS), Real-Time Messaging Protocol (RTMP), Dynamic Adaptive Streaming over HTTP (DASH), etc. In some embodiments, the server end 110 may include a local server or a cloud server. According to different service requirements, a local server corresponding to the area may be deployed in one or more areas. In some embodiments, the server end 110 may include a background processing server and a streaming media server. The background processing server can be used to analyze and process the video data. For example, the background processing server can perform audio and video encoding, transcoding, encryption, etc. on the original video data sent by the anchor end 120. The background processing server can also process the interactive data sent by the audience end 130. The streaming media server can transmit data based on the streaming media transmission protocol. In other embodiments, the server end 110 may also include a transcoding server, a storage server, an authentication server, etc. The transcoding server can be used to transcode the video data, the storage server can be used to cache and store the video data, and the authentication server can be used to verify the user's access rights. In some embodiments, the server end 110 can be a single computer device or a computing cluster composed of multiple computer devices, so as to provide more powerful computing power and more efficient response to user service requests.
[0051] The host terminal 120 can be used to generate video data (such as live video) in real time and perform streaming operations on the video data. The video data may include, for example, a video stream, a video file, or a data packet transmitted in blocks generated by real-time acquisition of live broadcast images by a camera device. In some embodiments, the host terminal 120 may be an electronic device such as a desktop computer, a smart phone, a laptop computer, and a tablet computer. In other embodiments, the host terminal 120 may also be a virtual computing instance. Exemplarily, the virtual computing instance may be a virtual computing resource running on a cloud server.
[0052] The viewer terminal 130 may include, but is not limited to, terminal devices such as desktop computers, smart phones, laptop computers, VR (Virtual Reality) devices, tablet computers, smart TVs, in-vehicle terminals, etc. The viewer terminal 130 may include a display screen and a processor, and the display screen may be used to present a graphical user interface. For example, the viewer terminal 130 may present a target video stream related to the live content sent by the server terminal 110 through the graphical user interface; for another example, the viewer terminal 130 may provide identifiers of interactive data through the graphical user interface, such as identifiers of interactive data for changing the style, special effects, character image, etc. of the live video. In some embodiments, the display screen may be separated from the human-computer interface device, and the user may operate on the graphical user interface through the human-computer interface device. The processor of the viewer terminal 130 may receive operation instructions generated by operating on the graphical user interface through the human-computer interface device, and the display screen may be used to present the graphical user interface. For example, the display screen may present a response result generated based on the operation instructions input through the graphical user interface. In other embodiments, the display screen may be a touch screen, and the touch screen may receive operation instructions input by the user based on the graphical user interface. When the user operates on the graphical user interface by touching the display screen, the graphical user interface may control the viewer terminal 130 to display local content in response to the received operation instructions, or may also control the viewer terminal 130 to display the content processed by the server terminal 110 based on the operation in response to the received operation instructions. For example, the operation instructions generated by the user acting on the graphical user interface may include instructions for selecting an interactive data identifier based on the graphical user interface. The viewer terminal processor may be configured to transmit the instructions for selecting the interactive data identifier to the server terminal 110 for processing after receiving them, and the server terminal 110 sends the processed content to the viewer terminal 130 for display.
[0053] The network 140 may be any form of wired or wireless network, or any combination thereof. As an example, the network 140 may be one or more combinations of a wired network, an optical fiber network, a telecommunication network, an internal network, the Internet, a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a public switched telephone network (PSTN), a Bluetooth network, etc. The network 140 may have multiple access points, and the server terminal 110, the host terminal 120, and the viewer terminal 130 may access the network 140 through the access points.
[0054] It should be noted that Figure 1 The schematic diagram of the live broadcast operation environment shown is only an example. The live broadcast operation environment described in the embodiments of this specification is for more clearly explaining the technical solutions of the embodiments of this specification, and does not constitute a limitation on the technical solutions provided by the embodiments of this specification. For example, Figure 1The quantities of the server side 110, the host side 120, and the audience side 130 in [it] are merely illustrative and are not used to limit the scope of patent protection of this application. According to the actual situation, there can be any number of server sides 110, host sides 120, and audience sides 130. As is known to those of ordinary skill in the art, with the development of live streaming technology and the emergence of new business scenarios, the technical solutions provided in the embodiments of this specification are equally applicable to similar technical problems.
[0055] In some related embodiments, the live video stream is usually the original video stream directly obtained by the video stream acquisition device (such as a camera device) on the host side, and the visual effect presentation is relatively single and lacks innovation. In some live streaming platforms, although the host can actively select some special effect elements (such as dynamic stickers worn on the host's head, virtual hanging ornaments for decorating the live streaming room, etc.) to add to the live streaming screen during the live broadcast, this requires the host's active operation, and the added special effect elements are fixed and cannot be dynamically adapted based on the content of the live streaming screen, lacking flexibility.
[0056] In view of this, some embodiments of this specification provide a live video generation method, which can obtain an original video stream, identify first information in the original video stream, convert the original video stream based on the first information, and generate a target video stream with a target comic style matching the original video stream. The first information can include at least one of the character action information, character expression information, or scene information in the original video stream. In this way, a comic-style target video stream matching the content of the original video stream can be automatically and dynamically generated, enriching the visual effect of the live streaming screen and enhancing the interest of the live streaming content. Moreover, through the automatic and dynamic generation method, manual intervention can be reduced, making the comic style in the generated target video stream more natural and vivid.
[0057] Figure 2 is an exemplary flowchart of a live video generation method shown according to some embodiments of this specification. Figure 2 The process 200 shown can be executed by a processing device. For example, it can be executed by Figure 1 the server side 110 shown. In some embodiments, the process 200 can be implemented by a live video generation device 1000 deployed on the processing device. As Figure 2 shown, in some embodiments, the process 200 can include the following steps.
[0058] Step 210, obtain an original video stream. In some embodiments, step 210 can be implemented by an acquisition module 1010.
[0059] In some embodiments, the original video stream may be a video stream sent from the live broadcast end. Exemplarily, the live broadcast end 120 may collect the live pictures of the host in real time through a video stream acquisition device (such as a camera device), and after encoding (such as using video encoding technologies such as H.264, HEVC, VVC, etc.) and compression, transmit it to the server end 110 in the form of a video stream. The server end 110 may correspondingly obtain the original video stream and perform subsequent processing.
[0060] In some embodiments, the original video stream may be a sequence of video frames with a time sequence, and the video frames may be static images. The pictures of the video frames may include people, objects, and backgrounds. Exemplarily, people may include, for example, the host, assistants, audience at the live broadcast site, etc.; objects may include, for example, tables, chairs, books, etc.; backgrounds may include, for example, indoor scenes, outdoor scenes, or natural landscapes, etc. Figure 3 is a schematic diagram of a video frame of an original video stream shown in some embodiments of this specification, such as Figure 3 shown, the picture of the video frame includes information such as the host, cosmetics, a tripod, and an indoor scene.
[0061] Step 220, identify the first information in the original video stream. In some embodiments, step 220 may be implemented by the identification module 1020.
[0062] In some embodiments, the first information may include at least one of the human action information, human expression information, or scene information in the original video stream. The following introduces these three types of information respectively.
[0063] In some embodiments, the human action information may include human key point information. Exemplarily, the human key points may include the head, shoulders, elbows, wrists, hips, knees, ankles, and various joint points of the hand, etc. Identifying the first information in the original video stream may include identifying the human key point information in the original video stream. Exemplarily, the human key point information in the original video stream may be identified through a pose estimation algorithm (such as the Open pose algorithm). For example, each video frame in the original video stream may be analyzed and processed through a model based on the pose estimation algorithm to obtain the human key point coordinates and the corresponding confidence levels. Exemplarily, the human key point coordinates may be the two-dimensional coordinates of each human key point in each video frame. The confidence level may be used to represent the reliability of the recognition result of the human key point. The higher the confidence level, the more reliable the recognition result. In some embodiments, the human key points may also be filtered according to the confidence level. For example, the human key points with a confidence level lower than a preset threshold may be filtered out to reduce the probability of the recognition result being incorrect and avoid affecting the subsequent processing process.
[0064] In some embodiments, the human action information may further include the human action category. Exemplarily, the human action category may include waving, jumping, turning in circles, bowing, turning around, playing the piano, singing, dance movements (such as floor movements in hip-hop), fitness movements (such as squats, planks, running), etc. Identifying the first information in the original video stream may include identifying the human key point information in the original video stream, and determining the human action category in the original video stream based on the human key point information. For the description of how to identify the human key point information, reference may be made to the description in the relevant embodiments of this step, which will not be elaborated here. In some embodiments, the human action category may be determined by the relative positions and movement trajectories of the human key point information. The relative positions of the human key point information may include, for example, the distances between the human key points, the angles formed by the connections of some human key points, etc. The movement trajectory of the human key point information may be determined based on the change of the coordinate values of the same human key point in a series of consecutive video frames. Further, a pre-trained action classification model (such as an action classification model implemented based on a convolutional neural network) may be used to analyze the relative positions of the human key point information or the movement trajectory of the human key point information to determine the human action category. In some embodiments, when determining the human action category, the human key point information may also be weighted based on the confidence level to determine the weight information of each human key point information. The higher the confidence level of the human key point information, the greater the influence on determining the human action category. In this way, the accuracy of determining the human action category can be avoided from being affected by the mis-identification of the human key point information.
[0065] In some embodiments, the human expression information may include the facial key point information. Exemplarily, the facial key points may include the forehead, eyebrows, eyes, nose, mouth, chin, ears, facial contour, etc. Identifying the first information in the original video stream may include identifying the facial key point information in the original video stream. Exemplarily, the facial key point information in the original video stream may be identified by a pose estimation algorithm (such as the Open pose algorithm) or a face recognition algorithm (such as Dlib, FaceNet). For example, each video frame in the original video stream may be analyzed and processed by a model based on a pose estimation algorithm or a face recognition algorithm to obtain the facial key point coordinates and the corresponding confidence levels. Exemplarily, the facial key point coordinates may be the two-dimensional coordinates of each facial key point in each video frame. The confidence level may be used to represent the reliability of the recognition result of the facial key point. The higher the confidence level, the more reliable the recognition result. In some embodiments, the facial key points may also be filtered according to the confidence level. For example, the facial key points with a confidence level lower than a preset threshold may be filtered out to reduce the probability of the recognition result being incorrect and avoid affecting the subsequent processing process.
[0066] In some embodiments, the facial expression information may further include a facial feature vector. The facial feature vector may be a high-dimensional vector obtained by extracting facial key points or facial features. Identifying the first information in the original video stream may include identifying the facial feature vector in the original video stream. Exemplarily, the facial feature vector in the original video stream may be identified by a facial recognition algorithm (such as Dlib, FaceNet). For example, the model of the facial recognition algorithm may be used to analyze and process each video frame in the original video stream, and output the facial feature vector (or simultaneously output the facial key point coordinates and the facial feature vector) and the corresponding confidence level. The confidence level may be used to represent the reliability of the recognition result of the facial feature vector. The higher the confidence level, the more reliable the recognition result. In some embodiments, the facial feature vector may also be filtered according to the confidence level. For example, the facial feature vectors with a confidence level lower than a preset threshold may be filtered out to reduce the probability of errors in the recognition result and avoid affecting the subsequent processing process.
[0067] In some embodiments, the human facial expression information may further include the human facial expression category. Exemplarily, the human facial expression category may include smiling, laughing, angry, rage, surprised, frowning, pain, etc. Identifying the first information in the original video stream may include identifying the facial key point information and / or facial feature vectors in the original video stream, and determining the human facial expression category in the original video stream based on the facial key point information and / or facial feature vectors. The description of how to identify the facial key point information and / or facial feature vectors may refer to the description in the relevant embodiments of this step, which will not be elaborated herein. In some embodiments, the human facial expression category may be determined by the relative positions and motion trajectories of the facial key point information. The relative positions of the facial key point information may include, for example, the distances between the facial key points, the angles formed by the connecting lines of some facial key points, etc. The motion trajectory of the facial key point information may be determined based on the coordinate values of the same facial key point in a plurality of consecutive video frames. Further, the facial key point information may be classified by a pre-trained expression classification model (such as an expression classification model implemented based on a convolutional neural network) to output the specific human facial expression category. In some other embodiments, the facial feature vectors may also be matched with the preset expression feature vectors in the expression feature library, and the human facial expression category may be determined according to the matching results. For example, if the similarity between the facial feature vector and the preset expression feature vector representing "rage" in the expression feature library is 60%, and the similarity between the facial feature vector and the preset expression feature vector representing "angry" in the expression feature library is 80%, the expression category corresponding to the preset expression feature vector with the highest similarity in the matching results may be determined as the human facial expression category corresponding to the facial feature vector. In some other embodiments, a classifier (such as a classifier based on a support vector machine or a neural network) may also be used to classify the facial feature vectors to output the corresponding human facial expression category. In some embodiments, when determining the human facial expression category, the facial key point information and / or facial feature vectors may also be weighted based on the confidence level to determine the weight information of each facial key point information and / or facial feature vector. The higher the confidence level of the facial key point information and / or facial feature vector, the greater the influence on determining the human facial expression category. In this way, the accuracy of determining the human facial expression category can be avoided from being affected by the misidentification of the facial key point information and / or facial feature vectors.
[0068] In some embodiments, the scene information may include at least one of person information, object information, and background information. Exemplarily, the person information may include person location and person category, the object information may include object location and object category, and the background information may include background area and background category. The person location may include, for example, human key point information, the object location may include, for example, the bounding box information of the object, and the background area may be determined based on the mask image of the background. The person category may include, for example, an anchor, an audience, etc., the object category may include, for example, a table, a chair, a book, etc., and the background category may include, for example, an indoor scene, an outdoor scene, a natural landscape, etc.
[0069] Exemplarily, for the recognition of the person location, the human key point information in the original video stream can be recognized through a pose estimation algorithm (such as the Open pose algorithm). The human key point information may be, for example, the human key point coordinates of each human key point in each video frame, such as the two-dimensional coordinates of each human key point in the original video frame. The pose estimation algorithm may further output the confidence corresponding to each human key point to represent the reliability of the recognition result of the human key point. The higher the confidence, the more reliable the recognition result. For the recognition of the person category, taking the recognition of the audience and the anchor in the original video stream as an example, it can be determined whether the person location is in the central area of the video frame based on the human key point coordinates. If so, the person category can be determined as the anchor, and if not, the person category can be determined as the audience. For another example, when there is only one person in the video frame, the person category can be determined as the anchor. For the recognition of the object location and object category, the bounding box information and category label of the object in the original video stream can be recognized through an object detection model (such as YOLO, Faster R-CNN). The bounding box information of the object is used as the object location, and the category label of the object is used as the object category. The object detection model may further output a confidence value of the object detection to represent the reliability of the recognition of the object location and object category. For the recognition of the background area and background category, the mask image and background category of the background in each video frame of the original video stream can be recognized through an image segmentation model (such as U-Net, DeepLab). For example, the classification label of each pixel in the video frame can be determined through the image segmentation model to represent which background category the pixel belongs to, such as indoor, outdoor, etc., and the position of the background area can be determined based on the mask image (such as a binary mask image) output by the image segmentation model.
[0070] Step 230, convert the original video stream based on the first information to generate a target video stream with a target comic style matching the original video stream. In some embodiments, step 230 may be implemented by the conversion module 1030.
[0071] In some embodiments, the target comic style may be a specific artistic style presented in the form of a comic for representing the overall visual effect of an image. For example, Japanese comic style, European and American comic style, Chinese-style comic style, ink painting style; for another example, cartoon style, hand-drawn comic style, retro comic style; for yet another example, a comic style suitable for presenting a sports state, a comic style suitable for presenting a fitness state, a comic style suitable for presenting a teaching state, etc. In some embodiments, the target video stream of the target comic style may include comic drawing information corresponding to the target comic style, and the comic drawing information may include color drawing information, line drawing information, comic character style, background style, etc. Exemplarily, for the Japanese comic style, the colors in the target video stream are relatively bright and vivid, the lines are delicate and smooth, the comic characters are Japanese comic characters, and the background is a Japanese comic background. For the European and American comic style, the color saturation in the target video stream is relatively low, and darker tones are mostly used, the lines are thick and heavy, the comic characters are European and American-style comic characters, and the background is a European and American-style comic background.
[0072] In some embodiments, the conversion of the original video stream based on the first information can be achieved through a pre-trained comic style conversion model. Among them, the comic style conversion model can be implemented based on a generative adversarial network (such as CycleGAN). The model architecture of the comic style conversion model may include a generator and a discriminator. Exemplarily, the generator can adopt a deep convolutional neural network (CNN) architecture such as U-Net or ResNet to receive input data and output a target image. The discriminator can adopt a standard convolutional neural network (CNN) architecture, which usually includes multiple convolutional layers and pooling layers, and is used to discriminate the probability that the image input to the discriminator is a real image (such as an image in a sample training set). By using a generative adversarial network to implement the comic style conversion model, high-quality images can be output, ensuring the real-time performance and low latency rate of the output images.
[0073] In some embodiments, the comic style conversion model can be obtained by training based on a sample training set, which can include video frame samples, first sample information extracted from the video frame samples, and comic style images matching the first sample information. In some embodiments, the video frame samples can be image frames extracted from video stream samples, such as image frames extracted from historical video streams of live broadcasts, and adjacent video frame samples can have temporal continuity. In some embodiments, the first sample information can include at least one of action sample information, expression sample information, and scene sample information. The content in the first sample information can correspond to the content in the first information. For example, if the first information includes human action information, the first sample information can correspondingly include action sample information; if the first information includes human expression information, the first sample information can correspondingly include expression sample information; if the first information includes scene information, the first sample information can correspondingly include scene sample information. The action sample information is similar to human action information, the expression sample information is similar to human expression information, and the scene sample information is similar to scene information. The difference is that the action sample information, expression sample information, and scene sample information are extracted from video stream samples (such as historical video streams of live broadcasts), while the human action information, human expression information, and scene information are extracted from the original video stream sent by the anchor side. For more descriptions of the action sample information, expression sample information, and scene sample information, reference can be made to the descriptions of the human action information, human expression information, and scene information in step 220 above, which will not be elaborated here. The method of extracting the first sample information from the video frame samples of the video stream sample is similar to the method of identifying the first information in the original video stream, and specific reference can be made to the description in step 220 above, which will not be elaborated here.
[0074] In some embodiments, the comic style images in the sample training set can be pre-drawn comic style images that match the first sample information in the video frame samples. Exemplarily, the first sample information extracted from the video frame samples can include human key point information, human action category (such as running), facial key point information, human expression category (such as concentration), human position, human category (such as the anchor), background area, background category (such as outdoor), and the corresponding comic style image can be a comic style image suitable for presenting the motion state. For example, the comic style image can include speed lines in comic form indicating the motion direction and speed, and drawn around the human position; the comic style image can also include an outdoor comic style background similar to the background in the video frame sample; the human in the comic style image can be a comic character related to the motion element.
[0075] In some embodiments, the training process of the comic style conversion model may include: After preprocessing the video frame samples in the sample training set and the first sample information extracted from the video frame samples (such as size adjustment, normalization, data augmentation, etc.), the input data is input into the generator of the comic style conversion model. The generator converts the input data into a comic style image that matches the first sample information. The discriminator determines the probability that the image generated by the generator is a real image (such as the comic style image corresponding to the input data of the generator in the sample training set). During the training process, the gradients are calculated by backpropagation through a loss function (such as adversarial loss, cycle consistency loss, content loss, pose consistency loss, expression consistency loss, etc.) to optimize the model parameters of the generator and the discriminator. After multiple iterative trainings and model parameter optimizations, a pre-trained comic style conversion model is obtained. Among them, the adversarial loss function is used to train the generator and the discriminator so that the image generated by the generator is as close as possible to the desired comic style image; the cycle consistency loss function is used to ensure that the image generated by the generator can be restored to the original video frame sample, maintaining content consistency; the content loss function is used to ensure that the content of the image generated by the generator is consistent with the original video frame sample; the pose consistency loss function and the expression consistency loss function are used to further constrain the actions and expressions of the image generated by the generator to be synchronized with the actions and expressions of the anchor in the original video frame sample.
[0076] In some embodiments, the pre-trained comic style conversion model may generate a target video stream with a target comic style that matches the original video stream based on the first information and each original video frame in the original video stream. Exemplarily, each original video frame in the original video stream and the first information extracted from each original video frame may be input into the pre-trained comic style conversion model. The generator of the comic style conversion model generates a target video frame with a target comic style that matches the original video frame according to the input data. Each target video frame is arranged in the time order corresponding to each original video frame to form a target video stream.
[0077] Exemplarily, for Figure 3For the original video frame shown, the first information identified from the original video frame may include: human key point information, human action category (such as introducing an item), facial key point information, human expression category (such as happy), human position, human category (such as a host), object position, object category (such as Japanese cosmetics, eyeshadow palette, and mirror, etc.), background area, background category (such as a Japanese-style decorated background). The original video frames in the original video stream and the first information extracted from each original video frame can be input into a pre-trained manga style conversion model. The manga style conversion model can generate a target video frame with a Japanese manga style matching the original video frame according to the input data. The target video frames are arranged in the time sequence corresponding to the original video frames to form a target video stream. Figure 4 is a schematic diagram of a video frame of a target video stream shown according to some embodiments of the present specification, as Figure 4 shown, the picture in the video frame of the target video stream can present a Japanese manga style. The host in the original video frame is replaced by a Japanese manga character, the background in the video frame of the target video stream is replaced by a Japanese-style background, and the actions and / or expressions of the manga character in the video frame of the target video stream are synchronized with the actions and / or expressions of the host in the original video stream.
[0078] In some other embodiments, the pre-trained manga style conversion model can also generate a target video stream with a target manga style matching the original video stream based on the fusion feature map or fusion feature vector generated from the first information and each original video frame of the original video stream. Exemplarily, a feature image can be generated based on the first information and fused with the original video frame to form a fusion feature map containing multi-channel inputs, and the fusion feature map is input into the pre-trained manga style conversion model; or for another example, feature extraction can be performed on the first information to obtain a corresponding feature vector, which is fused with the original video frame to form a fusion feature vector containing multi-channel inputs, and the fusion feature vector is input into the pre-trained manga style conversion model. Further, the manga style conversion model generates a target video frame with a target manga style matching the original video frame according to the input fusion feature map or fusion feature vector. The target video frames are arranged in the time sequence corresponding to the original video frames to form a target video stream. By fusing the first information with the original video frame, input data containing rich information can be generated, enabling the manga style conversion model to better understand and convert the original video frame and improving the conversion effect.
[0079] In some embodiments, the target video stream can also be compressed through video coding technologies (such as H.264, HEVC, VVC, etc.) to improve the efficiency of subsequent transmission of the target video stream. Further, video coding can be optimized. For example, coding parameters and algorithms can be optimized to reduce coding latency; the coding bit rate can also be dynamically adjusted according to the network bandwidth and video content to ensure a high-quality target video stream under different network conditions; multi-threading and multi-core processors can be utilized to encode the target video stream to improve the encoding speed; forward error correction codes can also be added during the encoding process to improve the reliability of the target video stream transmission; a packet loss recovery mechanism can also be designed to reduce the impact of packet loss in network transmission on the quality of the target video stream. By optimizing video coding, the efficiency of transmitting the target video stream can be effectively improved and the latency can be reduced.
[0080] In some embodiments, the first information can be information related to the host. Exemplarily, the human action information can include the action information of the host, the human expression information can include the expression information of the host, and the scene information can include the scene information of the live broadcast room where the host is located. For example, when there are multiple people in the video frame of the original video stream, the action information of the host, the expression information of the host, and the scene information of the live broadcast room where the host is located can be recognized, while ignoring the action information and expression information of other people (such as passers-by or viewers passing by in the scene of the live broadcast room) other than the host. In this way, when converting the original video stream based on the first information, a target video stream with a target comic style matching the original video stream can be generated based on the first information related to the host, which can avoid interference from non-host personnel to the live broadcast screen, reduce the consumption of computing resources, and improve the processing efficiency.
[0081] In some embodiments, a comic special effect matching the first information may also be added to the target video stream based on the first information. The comic special effect can be used to present the visual effects of characters, objects, or environments in the form of comics. Exemplarily, the comic special effect may include special effects related to actions presented in the form of comics, such as musical notes when playing the piano, line special effects for showing speed when running, etc. The comic special effect may also include special effects related to expressions presented in the form of comics, such as a smiling face special effect when smiling, an angry special effect when angry, etc. The comic special effect may also include special effects related to scenes presented in the form of comics, such as adding a glitter special effect, a flame special effect, etc. to the background area to create a specific atmosphere. The first information may include at least one of the character action information, character expression information, or scene information in the original video stream. For more descriptions of the first information, reference may be made to the description in step 220 above, which will not be elaborated here. In some embodiments, the comic special effect matching the first information may include the comic special effect matching the character action information, character expression information, or scene information of the first information. Exemplarily, for the comic special effect "line effect", it may match the character action information such as "waving", "running", etc.; for the comic special effect "angry face", it may match the character expression information such as "angry", "furious", etc.; for the comic special effect "flame effect", it may match the scene information such as "flame", etc., or may also match the character expression information such as "angry", "furious", etc.
[0082] In some embodiments, the first information may be matched with the preset special effects in the preset special effect library, and the comic special effect matching the first information may be determined according to the matching result. Exemplarily, when the character action category in the first information is waving, the comic special effect corresponding to the waving action may be matched in the preset special effect library, and according to the human key point information in the first information, the matched comic special effect may be added to the hand position of the character in the target video stream. For example, if the comic special effect "line effect" is matched with the character action category "waving" in the first information, a dynamic comic line effect may be added to the host's hand to enhance the visual effect of waving. When the character expression category in the first information is angry, the comic special effect with the expression element corresponding to the angry expression may be matched in the preset special effect library, and according to the facial key point information or the background area in the first information, the matched comic special effect may be added to the target video. For example, if the comic special effects "angry face" and "flame effect" are matched with the character expression category "angry" in the first information, a comic-style anger symbol may be added to the host's face, or a comic-style flame effect may be added to the background area to create an angry atmosphere. By adding the comic special effect matching the first information to the target video stream, the visual effect is further enhanced on the basis of the target video stream in the target comic style, the atmosphere of the live broadcast room is created, and the interest of the live broadcast content is further improved.
[0083] In some embodiments, it is also possible to determine whether to add a comic special effect that matches the first information to the target video stream based on a preset condition. In some embodiments, the preset condition may include that the time interval for adding the comic special effect is greater than a preset time interval. For example, the preset time interval is ten minutes. After adding the comic special effect "angry face" to the target video stream, no other comic special effects will be added within ten minutes, thereby reducing the frequent switching of the comic effect and avoiding discomfort to the user caused by the frequent replacement of the comic effect. In other embodiments, the preset condition may further include that the priority of the comic special effect to be added is higher than the priority of the comic special effect being displayed in the target video stream. For example, the comic special effect being displayed in the target video stream is "angry face". At this time, it is recognized that the action category of the person in the first information of the original video stream is "waving", and the comic special effect "line effect" corresponding to the action category of waving is matched from the preset special effect library. If the priority of the comic special effect "line effect" is higher than the priority of the comic special effect "angry face", then the comic special effect "line effect" is added to the target video stream. If the priority of the comic special effect "line effect" is lower than the priority of the comic special effect "angry face", then the comic special effect "angry face" continues to be displayed in the target video stream, thereby reducing the frequent switching of the comic effect and avoiding discomfort to the user caused by the frequent replacement of the comic effect. In some embodiments, when there are multiple comic special effects that match the first information, the comic special effect to be added to the target video stream can also be determined according to the priority of the comic special effect. For example, it is recognized that the facial expression category of the person in the first information of the original video stream is "angry", and the comic special effects corresponding to anger include "angry face" and "flame". The priority of the comic special effect "angry face" is higher than the priority of the comic special effect "flame", then the comic special effect "angry face" is determined as the comic special effect to be added to the target video stream, so as to limit the number of comic special effects presented in the target video stream and avoid affecting the visual effect of the comic special effect due to too many comic special effects. In some embodiments, a smooth transition effect can also be added between adjacent comic special effects to present a natural visual effect when switching comic special effects.
[0084] To enhance the interactivity and sense of participation of the audience watching the live broadcast, candidate comic elements can also be provided for the audience to select, and the target video stream can be updated according to the audience's selection. Figure 5 It is an exemplary flowchart of another live video generation method shown according to some embodiments of this specification. Figure 5 The process 500 shown can be executed by a processing device. For example, it can be executed by Figure 1 the server side 110 shown. In some embodiments, the process 500 can be a subsequent step of the process 200. In some embodiments, the process 500 can be implemented by the target video stream update module 1040 or the conversion module 1030 in the live video generation device 1000 deployed on the processing device. As Figure 5As shown, in some embodiments, process 500 may include the following steps.
[0085] Step 510, obtaining the selection information of candidate comic elements sent by the viewer terminal.
[0086] In some embodiments, the candidate comic elements may include at least one of candidate comic styles, candidate comic special effects, or candidate comic characters. Exemplarily, the candidate comic styles may include, for example, Japanese comic style, European and American comic style, Chinese-style comic style, ink painting style, cartoon style, hand-drawn comic style, retro comic style, comic style suitable for presenting a sports state, comic style suitable for presenting a fitness state, comic style suitable for presenting a teaching state, etc. The candidate comic special effects may include, for example, lightning, flame, bubbles, musical notes when playing the piano, speed line special effects for running, smiling face special effects, angry special effects, etc. The candidate comic characters may include, for example, Japanese comic characters, European and American comic characters, Chinese-style comic characters, ink painting style comic characters, retro comic characters, etc.
[0087] In some embodiments, the viewer terminal 130 may provide an interactive interface, which may be, for example, an interactive area displayed in the graphical user interface of the viewer terminal 130. The identification of candidate comic elements (such as text, icons, thumbnails, etc.) may be displayed in the interactive interface. The user may select the identification of the candidate comic elements in the interactive interface. The viewer terminal 130 may respond to the user's selection operation and send the candidate comic elements corresponding to the selection operation to the server terminal 110. The server terminal 110 obtains the selection information of the candidate comic elements sent by each viewer terminal. In some embodiments, the viewer terminal 130 may receive the candidate comic elements sent by the server terminal 110 and display the identification of the candidate comic elements in the interactive interface. In some other embodiments, the viewer terminal 130 may also retrieve the candidate comic elements from the local stored data and display the identification of the candidate comic elements in the interactive interface.
[0088] Step 520, determining the updated target comic elements based on the selection information of the candidate comic elements, and generating an updated target video stream based on the updated target comic elements.
[0089] In some embodiments, the updated target comic elements may be determined based on the statistical quantity of the selection information of the candidate comic elements, and an updated target video stream may be generated based on the updated target comic elements. Further, the updated target video stream may be sent to each viewer terminal. Different viewer terminals may correspond to the same updated target video stream, and the same updated target video stream may be displayed through the interactive interface, so as to ensure the consistency of the live broadcast effect presented by each viewer terminal.
[0090] In some embodiments, the candidate comic element with the largest number of selections can be determined as the updated target comic element based on the statistical quantity of the selection information of the candidate comic elements. Exemplarily, assuming that in the selection information of the candidate comic elements, the number of selections for the candidate comic style of the European and American comic style is 20 votes, the number of selections for the Japanese comic style is 30 votes, and the number of selections for the Chinese-style comic style is 50 votes, then the Chinese-style comic style with the largest number of selections can be determined as the updated target comic style. In some embodiments, the statistical quantity of the selection information of the candidate comic elements can also be updated periodically, such as updating the statistical quantity of the selection information of the candidate comic elements every X seconds (or every half hour), and determining the candidate comic element with the largest number of selections as the updated target comic element.
[0091] In some embodiments, the weight information of each candidate comic element can be determined based on the statistical quantity of the selection information of the candidate comic elements, and a comprehensive comic element can be generated based on the weight information as the updated target comic element. Among them, the comprehensive comic element integrates the characteristics of each candidate comic element. Exemplarily, assuming that in the selection information of the candidate comic elements, the number of selections for the candidate comic style of the European and American comic style is 20 votes, the number of selections for the Japanese comic style is 30 votes, and the number of selections for the Chinese-style comic style is 50 votes, then the weight information of the three candidate comic styles is 20 / 100 = 0.2, 30 / 100 = 0.3, and 50 / 100 = 0.5 respectively. In the corresponding generated comprehensive comic style, the Chinese-style comic style accounts for 50%, the Japanese comic style accounts for 30%, and the European and American comic style accounts for 20%. In this way, the generated comprehensive comic style not only includes each comic style but also reflects the voting preferences of the audience, thus providing a richer and more personalized visual effect. Compared with a single comic style, the comprehensive comic style can better meet the aesthetic needs of different audiences and present a comic style effect with mixed aesthetic characteristics.
[0092] In some other embodiments, the updated target comic element corresponding to each client can also be determined based on the selection information of the candidate comic elements of each client, and an updated target video stream corresponding to each client can be generated based on each updated target comic element. That is to say, different clients can correspond to different updated target video streams. Further, each updated target video stream can be sent to the corresponding client respectively, so that each client can display the personalized updated target video stream through the interaction interface.
[0093] In some embodiments, when the selection information of the candidate comic elements includes the candidate comic style, the updated target comic style can be determined based on the selection information of the candidate comic style, and the original video stream can be converted based on the updated target comic style to generate an updated target video stream, and the updated target video stream has the updated target comic style.
[0094] In some embodiments, based on the updated target comic style, the original video stream is transformed to generate an updated target video stream, which can be achieved through a pre-trained comic style transformation model. The comic style transformation model can be obtained by training based on a sample training set, which can include video frame samples, first sample information extracted from the video frame samples, comic style labels, and comic images with a comic style corresponding to the comic style labels that match the video frame samples. Exemplarily, the comic style labels can include Japanese, European and American, Chinese style, comprehensive, ink painting, cartoon, hand-drawn, retro, sports, fitness, teaching, etc. The descriptions of the video frame samples and the first sample information can refer to the descriptions in step 230 above and will not be elaborated here.
[0095] In some embodiments, the training process of the comic style transformation model can include: after preprocessing the video frame samples, the first sample information extracted from the video frame samples, and the comic style labels in the sample training set (such as resizing, normalization, data augmentation, etc.), they are used as input data and input into the generator of the comic style transformation model. Through the generator, the input data is transformed into a comic image with a comic style corresponding to the comic style label that matches the video frame sample. The discriminator is used to determine the probability that the image generated by the generator is a real image (such as the comic image with the comic style corresponding to the comic style label in the sample training set). During the training process, the gradient is calculated by backpropagation through a loss function (such as adversarial loss, cycle consistency loss, content loss, pose consistency loss, expression consistency loss, etc.) to optimize the model parameters of the generator and the discriminator. After multiple iterations of training and model parameter optimization, a pre-trained comic style transformation model is obtained. In some embodiments, during the training process, comic images of multiple comic styles can also be used to train the comic style transformation model so that the generator can migrate and switch between different comic styles.
[0096] Figure 6 It is a schematic diagram of a video frame of an updated target video stream according to some embodiments of this specification. Exemplarily, when the statistical quantity of the selection information of the candidate comic elements shows that the number of European and American style comic styles is the largest, the updated target comic style can be determined as the European and American style comic style. Further, the original video frames in the original video stream (such as Figure 3 shown), the first information extracted from each original video frame, and the European and American style comic style are input into the pre-trained comic style transformation model. The generator of the comic style transformation model generates target video frames in the European and American style comic style as shown in Figure 6 shown. The target video frames are arranged in the time sequence corresponding to each original video frame to form an updated target video stream. In this way, the migration of the comic style can be realized based on the audience's selection, increasing the audience's sense of participation and enhancing the interactivity.
[0097] The comic style conversion model can convert the original video stream based on the first information to generate a target video stream with a target comic style that matches the original video stream, or can also convert the original video stream based on an updated target comic style (the updated target comic style can be determined based on the selection information of candidate comic elements) to generate an updated target video stream. Here, the same comic style conversion model is used for data conversion, but different training processes are involved. The specific training process can refer to the descriptions in the aforementioned step 230 and step 520. Using the same comic style conversion model but different training data and training processes can enable the comic style conversion model to perform different processes according to different input data. The input data can include a specified comic style (such as the updated target comic style determined based on the selection information of candidate comic elements), or can not specify a comic style. When the comic style is not specified in the input data, the comic style conversion model can generate a target video frame with a target comic style that matches the original video frame based on the input data (such as the original video frame and the first information). When the input data includes a specified comic style (such as the updated target comic style determined based on the selection information of candidate comic elements), the comic style conversion model can generate a video frame with the updated target comic style based on the input data (such as the original video frame, the first information, and the updated target comic style), thereby realizing the style transfer.
[0098] In some embodiments, model compression and acceleration techniques, such as pruning, quantization, knowledge distillation, etc., can also be adopted during the training process of the comic style conversion model to improve the inference speed of the generator and enhance the conversion efficiency of the comic style conversion model. Further, a low-latency generator and discriminator architecture can also be used to reduce the computational overhead of the comic style conversion model and enhance the real-time performance of the conversion of the comic style conversion model.
[0099] In some embodiments, when the selection information of candidate comic elements includes candidate comic special effects or candidate comic characters, the target comic special effects or target comic characters can be determined based on the selection information of the candidate comic special effects or candidate comic characters. Further, based on the target comic special effects or target comic characters, the target video stream with the target comic style is adjusted to generate an updated target video stream, and the updated target video stream can include the target comic special effects or target comic characters.
[0100] In some embodiments, target comic special effects can be added to the target video stream with the target comic style. Exemplarily, assuming the target comic special effect is a flame special effect, then a flame special effect layer can be superimposed on the target video stream to generate an updated target video stream, and the updated target video stream can include the flame special effect. Among them, the special effect layer can be generated through image processing techniques (such as OpenCV).
[0101] In some embodiments, the comic special effects in the target video stream with the target comic style can be replaced with the target comic special effects. Exemplarily, assuming that the target comic special effect is a flame special effect, when a lightning special effect has been superimposed on the target video stream, the lightning special effect can be replaced with the flame special effect to generate an updated target video stream, and the updated target video stream can include the flame special effect.
[0102] In some embodiments, the characters in the target video stream with the target comic style can be replaced with the target comic characters. Exemplarily, assuming that the target comic character is a Japanese comic character, the host can be segmented from the target video stream through image segmentation techniques (such as the image segmentation model U-Net, DeepLab, etc.) or object detection techniques (such as the object detection model YOLO, Faster R-CNN) and replaced with the character image of the Japanese comic character.
[0103] In some embodiments, candidate comic elements can also be determined based on preset information or historical selection information of the viewer terminal for the viewer terminal to select based on the interaction interface. Exemplarily, one or more preset comic styles, one or more preset comic special effects, or one or more preset comic characters can be used as candidate comic elements for the viewer terminal to select based on the interaction interface. Exemplarily, the candidate comic elements can also be determined according to the ranking order of the number of times the historical candidate comic elements are selected in the historical selection information. For example, the top 10 historical candidate comic elements ranked by the number of times they are selected are used as candidate comic elements for the viewer terminal to select based on the interaction interface.
[0104] In some embodiments, the display order of the candidate comic elements can also be updated based on the real-time selection information of the viewer terminal for the candidate comic elements. Exemplarily, based on the real-time selection information of the viewer terminal for the candidate comic elements, the selection quantity of each candidate comic element can be determined, and each candidate comic element can be sorted according to the selection quantity of the candidate comic elements, so as to guide the user of the viewer terminal to select popular options and increase the interest of interaction.
[0105] In one or more embodiments of this specification, by determining updated target comic elements based on the selection information of the candidate comic elements and generating an updated target video stream based on the updated target comic elements, an audience interaction mechanism is provided to update the target video stream according to the audience's selection, enhancing the audience's sense of participation and interactivity during the live broadcast. The audience can select and switch different candidate comic elements according to their own preferences, rather than being limited to a fixed comic style, comic special effect, or comic character.
[0106] In some embodiments, the target video stream generated by the server side 110 or the updated target video stream generated based on the selection information of candidate comic elements can be presented to the user through the viewer side 130 for the user to watch the live broadcast. The viewer side 130 can transmit data with the server side 110 and present the video data (such as the target video stream or the updated target video stream) sent by the server side 110 through a graphical user interface, so as to present a live broadcast picture for the user. This specification also provides a method for displaying a live broadcast picture. Figure 7 It is an exemplary flowchart of a method for displaying a live broadcast picture according to some embodiments of this specification. Figure 7 The process 700 shown can be executed by a terminal device. For example, it can be executed by Figure 1 the viewer side 130 shown. In some embodiments, the process 700 can be implemented by a live broadcast picture display device 1100 deployed on the terminal device. As Figure 7 shown, in some embodiments, the process 700 can include the following steps.
[0107] Step 710, receiving a target video stream with a target comic style sent by the server side. In some embodiments, step 710 can be implemented by a receiving module 1110 in the live broadcast picture display device 1100.
[0108] In some embodiments, the target video stream can be generated by the server side based on the first information in the original video stream sent from the live broadcast end to convert the original video stream, and the target video stream has a target comic style matching the original video stream. In some embodiments, the first information can include at least one of the character action information, character expression information, or scene information in the original video stream. For more descriptions of the first information, reference can be made to the description in step 220 above, and details will not be elaborated here.
[0109] For the description of the target video stream with a target comic style, and how the server side converts the original video stream based on the first information in the original video stream to generate a target video stream with a target comic style matching the original video stream, reference can be made to the description in step 230 above, and details will not be elaborated here.
[0110] Step 720, displaying the target video stream with a target comic style in the live broadcast picture. In some embodiments, step 720 can be implemented by a display module 1120 in the live broadcast picture display device 1100.
[0111] In some embodiments, the viewer side 130 can provide a graphical user interface, and the live broadcast picture can be displayed in the graphical user interface. The viewer side 130 can obtain the target video stream sent by the server side 110 in real time and display the target video stream with a target comic style in the live broadcast picture.
[0112] In one or more embodiments of the present specification, the target video stream is generated by the server side based on the first information in the original video stream sent by the live broadcast end, and the target video stream has a target comic style that matches the original video stream. The viewer side receives the target video stream with the target comic style sent by the server side and displays the target video stream with the target comic style in the live broadcast screen, so that the target video stream with a specific comic style that matches the content of the original video stream is presented in the live broadcast screen of the viewer side, making the visual effect of the live broadcast screen more rich and enhancing the interest of the live broadcast content.
[0113] In order to enhance the interactivity and participation of the viewers watching the live broadcast, candidate comic elements can also be provided on the viewer side (such as a WEB page, a mobile APP, etc.) for the viewers to select, and the target video stream can be updated according to the viewers' selections. Figure 8 It is an exemplary flowchart of another live broadcast screen display method shown in some embodiments of the present specification. Figure 8 The process 800 shown can be executed by a terminal device. For example, by Figure 1 the viewer side 130 shown. In some embodiments, the process 800 can be implemented by the target video stream update module 1130 in the 1100 of the live broadcast screen display device deployed on the terminal device. As Figure 8 shown, in some embodiments, the process 800 can include the following steps.
[0114] Step 810, display candidate comic elements.
[0115] In some embodiments, the viewer side 130 can provide an interaction interface, and the interaction interface can be, for example, an interaction area displayed in the graphical user interface of the viewer side 130. The interaction interface can include a Web-side interaction interface and / or a mobile-side interaction interface. For the Web-side interaction interface, exemplarily, the interaction interface can be created based on HTML, CSS, and JavaScript. For the mobile-side interaction interface (such as a mobile APP), exemplarily, the interaction interface can be created based on a mobile development framework (such as React Native, Flutter).
[0116] In some embodiments, the identifiers (such as text, icons, thumbnails, etc.) of the candidate comic elements can be displayed in the interaction interface of the viewer side 130 for the user to select. In some embodiments, the candidate comic elements can include at least one of candidate comic styles, candidate comic special effects, or candidate comic characters. For more descriptions of the candidate comic elements, reference can be made to the description in step 510 above, and details will not be elaborated here.
[0117] In some embodiments, the viewer terminal 130 can receive candidate comic elements sent by the server terminal 110 and display the identifiers of the candidate comic elements in the interaction interface. In other embodiments, the viewer terminal 130 can also retrieve candidate comic elements from local stored data and display the identifiers of the candidate comic elements in the interaction interface.
[0118] Step 820: Obtain the selection information of the candidate comic elements input based on the interaction interface and send it to the server terminal.
[0119] In some embodiments, the user can select the identifier of the candidate comic element displayed in the interaction interface of the viewer terminal 130. The viewer terminal 130 responds to the selection operation to obtain the selection information of the candidate comic element and sends the selection information of the candidate comic element to the server terminal. Exemplarily, various candidate comic styles can be displayed in the interaction interface of the viewer terminal 130, such as Japanese comic style, European and American comic style, Chinese-style comic style, ink painting style, etc. The user can select the European and American comic style in the interaction interface. The viewer terminal 130 responds to the selection operation to obtain the selection information of the European and American comic style and sends the selection information of the European and American comic style to the server terminal.
[0120] Step 830: Receive and display the updated target video stream sent by the server terminal.
[0121] In some embodiments, the live broadcast screen can be displayed in the graphical user interface of the viewer terminal 130. After the viewer terminal 130 receives the updated target video stream sent by the server terminal 110, it displays the updated target video stream in the live broadcast screen.
[0122] In some embodiments, the updated target video stream is generated by the server terminal based on the selection information of the candidate comic elements. For the description of the updated target video stream and how the server terminal generates the updated target video stream based on the selection information of the candidate comic elements, reference can be made to the description in step 520 above and will not be elaborated here.
[0123] In some embodiments, when the selection information of the candidate comic elements includes the candidate comic style, the updated target video stream has an updated target comic style determined based on the selection information of the candidate comic style. Exemplarily, the server terminal 110 can convert the original video stream based on the updated target comic style to generate the updated target video stream, so that the updated target video stream has the updated target comic style, and send the updated target video stream to the viewer terminal 130. For the specific description of this part, reference can be made to the description in step 520 above and will not be elaborated here.
[0124] In some other embodiments, when the selection information of the candidate comic elements includes candidate comic special effects or candidate comic characters, the updated target video stream includes the target comic special effects or target comic characters determined based on the selection information of the candidate comic special effects or candidate comic characters. Exemplarily, the server 110 may determine the target comic special effects or target comic characters based on the selection information of the candidate comic special effects or candidate comic characters, and adjust the target video stream with the target comic style based on the target comic special effects or target comic characters to generate an updated target video stream, and the updated target video stream may include the target comic special effects or target comic characters. Further, the updated target video stream is sent to the viewer terminal 130. For the specific description of this part, reference may be made to the description in step 520 above, which will not be elaborated here.
[0125] In some embodiments, the updated target video stream is generated by the server based on the updated target comic elements, and the updated target comic elements are the candidate comic elements with the largest number of selections or the comprehensive comic elements, and the comprehensive comic elements are generated based on the weight information of each candidate comic element. For the description of how the server generates the updated target video stream based on the updated target comic elements and how to determine the candidate comic elements with the largest number of selections or the comprehensive comic elements as the updated target comic elements, reference may be made to the description in step 520 above, which will not be elaborated here.
[0126] In some embodiments, the server 110 may perform data transmission with the anchor terminal 120. The original video stream obtained by the server 110 may be sent by the anchor terminal 120, and the target video stream generated by the server 110 or the updated target video stream generated based on the selection information of the candidate comic elements may be presented to the user through the anchor terminal 120. The anchor terminal 120 may present the target video stream or the updated target video stream sent by the server 110 through a graphical user interface, so as to present a live broadcast screen for the users of the anchor terminal. This specification also provides a method for displaying a live broadcast screen. Figure 9 It is an exemplary flowchart of a method for displaying a live broadcast screen according to some embodiments of this specification. Figure 9 The process 900 shown can be executed by a terminal device. For example, it can be executed by Figure 1 the anchor terminal 120 shown. In some embodiments, the process 900 can be implemented by a live broadcast screen display device 1200 deployed on the terminal device. As Figure 9 shown, in some embodiments, the process 900 may include the following steps.
[0127] Step 910, obtain and send the original video stream of the live broadcast screen to the server. In some embodiments, step 910 can be implemented by the obtaining module 1210 in the live broadcast screen display device 1200.
[0128] In some embodiments, the live video of the host can be collected in real time by a video stream acquisition device (such as a camera device). After encoding (such as using video encoding technologies such as H.264, HEVC, VVC, etc.) and compression, it is transmitted to the server side 110 in the form of a video stream. For more descriptions of the original video stream, reference can be made to the description in step 210 above, which will not be elaborated here.
[0129] Step 920: Receive the target video stream in the target comic style sent by the server side. In some embodiments, step 920 can be implemented by the receiving module 1220 in the live video display device 1200.
[0130] In some embodiments, the target video stream can be generated by the server side converting the original video stream based on the first information in the original video stream and has a target comic style matching the original video stream. In some embodiments, the first information can include at least one of the character action information, character expression information, or scene information in the original video stream. For more descriptions of the first information, reference can be made to the description in step 220 above, which will not be elaborated here. For the description of the target video stream in the target comic style and how the server side converts the original video stream based on the first information in the original video stream to generate the target video stream with a target comic style matching the original video stream, reference can be made to the description in step 230 above, which will not be elaborated here.
[0131] Step 930: Display the target video stream in the live video. In some embodiments, step 930 can be implemented by the display module 1230 in the live video display device 1200.
[0132] In some embodiments, the live video can be displayed in the graphical user interface provided by the host side 120, and the target video stream in the target comic style can be displayed in the live video.
[0133] In one or more embodiments of this specification, the target video stream can be generated by the server side converting the original video stream based on the first information in the original video stream and has a target comic style matching the original video stream. The host side obtains and sends the original video stream of the live video to the server side. Further, it receives the target video stream in the target comic style sent by the server side and displays the target video stream in the live video, so that a target video stream in a specific comic style matching the content of the original video stream is presented in the live video of the host side, making the visual effect of the live video more rich and enhancing the interest of the live content.
[0134] In some embodiments, the server side 110 may perform data transmission with the host side 120 and the viewer side 130. The server side 110 may obtain an original video stream from the host side 120, and generate a target video stream based on the content of the original video stream, or generate an updated target video stream based on the selection information of candidate comic elements sent by the viewer side 130. This specification also provides a live video generation device. Figure 10 It is an exemplary block diagram of a live video generation device shown according to some embodiments of this specification. In some embodiments, the live video generation device 1000 may be deployed on the server side 110. As Figure 10 shown, in some embodiments, the live video generation device 1000 may include an acquisition module 1010, an identification module 1020, and a conversion module 1030.
[0135] The acquisition module 1010 is configured to acquire an original video stream.
[0136] The identification module 1020 is configured to identify first information in the original video stream, where the first information includes at least one of human action information, human expression information, or scene information in the original video stream.
[0137] The conversion module 1030 is configured to convert the original video stream based on the first information to generate a target video stream with a target comic style matching the original video stream.
[0138] In some alternative embodiments, the live video generation device 1000 may further include a target video stream update module 1040, configured to acquire the selection information of candidate comic elements sent by the viewer side, where the candidate comic elements include at least one of a candidate comic style, a candidate comic special effect, or a candidate comic character; determine updated target comic elements based on the selection information of the candidate comic elements, and generate an updated target video stream based on the updated target comic elements. In some embodiments, the target video stream update module 1040 and the conversion module 1030 may be the same module, and the conversion module 1030 may also be configured to acquire the selection information of candidate comic elements sent by the viewer side, determine updated target comic elements based on the selection information of the candidate comic elements, and generate an updated target video stream based on the updated target comic elements.
[0139] In some alternative embodiments, when the selection information of the candidate comic elements includes a candidate comic style, the target video stream update module 1040 may further be configured to determine an updated target comic style based on the selection information of the candidate comic style; convert the original video stream based on the updated target comic style to generate an updated target video stream, and the updated target video stream has the updated target comic style.
[0140] In some optional embodiments, the target video stream updating module 1040 may further be configured to determine, based on the statistical quantity of the selection information of the candidate comic elements, the candidate comic element with the largest selection quantity as the updated target comic element; or determine the weight information of each candidate comic element based on the statistical quantity of the selection information of the candidate comic elements, and generate a comprehensive comic element as the updated target comic element based on the weight information, where the comprehensive comic element integrates the characteristics of each candidate comic element.
[0141] In some optional embodiments, when the selection information of the candidate comic elements includes candidate comic special effects or candidate comic characters, the target video stream updating module 1040 may further be configured to determine the target comic special effects or target comic characters based on the selection information of the candidate comic special effects or candidate comic characters; adjust the target video stream with the target comic style based on the target comic special effects or target comic characters to generate an updated target video stream, where the updated target video stream includes the target comic special effects or target comic characters.
[0142] In some optional embodiments, the target video stream updating module 1040 may further be configured to add target comic special effects to the target video stream with the target comic style; or replace the comic special effects in the target video stream with the target comic style with the target comic special effects; or replace the characters in the target video stream with the target comic style with the target comic characters.
[0143] In some optional embodiments, the live video generating device 1000 may further include a candidate comic element determining module 1050, configured to determine candidate comic elements based on preset information or historical selection information of the viewer terminal for the viewer terminal to select based on the interaction interface.
[0144] In some optional embodiments, the human motion information includes human body key point information and / or human motion categories; the recognition module 1020 may further be configured to recognize the human body key point information in the original video stream; determine the human motion categories in the original video stream based on the human body key point information.
[0145] In some optional embodiments, the human expression information includes at least one of human expression categories, facial key point information, and facial feature vectors; the recognition module 1020 may further be configured to recognize the facial key point information and / or facial feature vectors in the original video stream; determine the human expression categories in the original video stream based on the facial key point information and / or facial feature vectors.
[0146] In some alternative embodiments, the scene information includes at least one of person information, object information, and background information; the recognition module 1020 can also be used to recognize at least one of person information, object information, and background information in the original video stream; wherein, the person information includes person position and person category, the object information includes object position and object category, and the background information includes background area and background category; the person position includes human key point information; the object position includes the bounding box information of the object; the background area is determined based on the mask image of the background.
[0147] In some alternative embodiments, the live video generation device 1000 can also include a comic special effect adding module 1060, which is used to add comic special effects matching the first information to the target video stream based on the first information.
[0148] In some alternative embodiments, the conversion module 1030 can also be used to convert the original video stream based on the first information through a pre-trained comic style conversion model; the pre-trained comic style conversion model generates a target video stream with a target comic style matching the original video stream based on the first information and each original video frame in the original video stream, or based on the fused feature map or fused feature vector generated from the first information and each original video frame of the original video stream.
[0149] In some embodiments, the server side 110 can perform data transmission with the viewer side 130. The target video stream generated by the server side 110 or the updated target video stream generated based on the selection information of the candidate comic elements can be presented to the user through the viewer side 130. The viewer side 130 can present the target video stream or the updated target video stream sent by the server side 110 through a graphical user interface, so as to present a live broadcast screen to the user. This specification also provides a live broadcast screen display device. Figure 11 It is an exemplary block diagram of a live broadcast screen display device shown according to some embodiments of this specification. In some embodiments, the live broadcast screen display device 1100 can be deployed on the viewer side 130. As Figure 11 shown, in some embodiments, the live broadcast screen display device 1100 can include a receiving module 1110 and a display module 1120.
[0150] The receiving module 1110 is configured to receive the target video stream with a target comic style sent by the server side. The target video stream is generated by the server side based on the first information in the original video stream sent from the live broadcast end to convert the original video stream, and the target video stream has a target comic style matching the original video stream; wherein, the first information includes at least one of person action information, person expression information, or scene information in the original video stream.
[0151] The display module 1120 is configured to display the target video stream with a target comic style in the live broadcast screen.
[0152] In some alternative embodiments, the live video display device 1100 may further include a target video stream update module 1130, configured to display candidate comic elements, where the candidate comic elements include at least one of a candidate comic style, candidate comic special effects, or candidate comic characters; obtain selection information of the candidate comic elements input based on an interaction interface and send it to the server side; receive and display the updated target video stream sent by the server side, where the updated target video stream is generated by the server side based on the selection information of the candidate comic elements.
[0153] In some embodiments, the server side 110 may perform data transmission with the host side 120. The original video stream obtained by the server side 110 may be sent by the host side 120. The target video stream generated by the server side 110 or the updated target video stream generated based on the selection information of the candidate comic elements may be presented to the user through the host side 120. The host side 120 may present the target video stream or the updated target video stream sent by the server side 110 through a graphical user interface, so as to present a live video to the user of the host side. This specification also provides a live video display device. Figure 12 It is an exemplary block diagram of a live video display device shown according to some embodiments of this specification. In some embodiments, the live video display device 1200 may be deployed on the host side 120. As Figure 12 shown, in some embodiments, the live video display device 1200 may include an acquisition module 1210, a reception module 1220, and a display module 1230.
[0154] The acquisition module 1210 is configured to acquire and send the original video stream of the live video to the server side.
[0155] The reception module 1220 is configured to receive the target video stream with a target comic style sent by the server side, where the target video stream is generated by the server side based on the first information in the original video stream to convert the original video stream and has a target comic style matching the original video stream, where the first information includes at least one of the character action information, character expression information, or scene information in the original video stream.
[0156] The display module 1230 is configured to display the target video stream in the live video.
[0157] This specification also provides a live video system. Figure 13 It is an exemplary block diagram of a live video system shown according to some embodiments of this specification. As Figure 13 shown, in some embodiments, the live video system 1300 may include a server side 1310, a host side 1320, and an audience side 1330.
[0158] The server side 1310 is configured to obtain the original video stream sent by the autonomous live broadcast end, identify the first information in the original video stream, convert the original video stream based on the first information, generate a target video stream with a target comic style matching the original video stream, and send it to the live broadcast end and the viewer end. The first information includes at least one of the human action information, human expression information, or scene information in the original video stream.
[0159] The live broadcast end 1320 is configured to obtain and send the original video stream to the server side, and obtain and display the target video stream with the target comic style sent by the server side.
[0160] The viewer end 1330 is configured to obtain and display the target video stream with the target comic style sent by the server side.
[0161] For more content about each module, device, and system, reference can be made to Figures 2 - 9 the relevant description, which will not be elaborated here. It should be understood that Figures 10 - 13 the devices, systems, and their modules shown can be implemented in various ways. For example, in some embodiments, they can be implemented through hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art can understand that the above methods, devices, and systems can be implemented using computer-executable instructions and / or control codes included in a processor. For example, such codes are provided in carrier media such as disks, CDs, or DVD-ROMs, or in the memories of programmable devices. The devices and their modules in this specification can be implemented not only by hardware circuits such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or programmable hardware devices such as field programmable gate arrays and programmable logic devices, but also by software executed by various types of processors, or by a combination of the above hardware circuits and software (e.g., firmware).
[0162] It should be noted that the above descriptions of the devices, systems, and modules are only for convenience of description and do not limit this specification to the scope of the examples given. It can be understood that for those skilled in the art, after understanding the principle of the device, they can, without departing from this principle, arbitrarily combine the various modules to form a sub-device connected to other modules. Or split some modules to obtain more modules or multiple units under that module. Such deformations are all within the scope disclosed in this specification.
[0163] Some embodiments of this specification also provide a computer device, which includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the methods described in this specification can be implemented. Figures 2 - 9 The methods shown.
[0164] Some embodiments of this specification also provide a computer-readable storage medium, in which computer instructions are stored. When the computer instructions are executed by a processor, the methods described in this specification can be implemented. Figures 2 - 9 The methods shown.
[0165] Some embodiments of this specification also provide a computer program product, including a computer program. When at least part of the computer program is executed by a processor, the methods described in this specification can be implemented. Figures 2 - 9 The methods shown. In some embodiments, the computer program product may only relate to a computer program, which may be carried by a storage medium or a processing device. In other embodiments, the computer program product may also be a storage medium or a processing device containing the aforementioned computer program. The processing device may include one or more processors and a storage medium.
[0166] The beneficial effects that may be brought about by the embodiments of this specification include, but are not limited to: (1) By obtaining the original video stream, identifying the first information in the original video stream, converting the original video stream based on the first information, and generating a target video stream with a target comic style matching the original video stream, the first information may include at least one of the character action information, character expression information, or scene information in the original video stream. In this way, a target video stream with a comic style matching the content of the original video stream can be automatically and dynamically generated, enriching the visual effect of the live broadcast screen and enhancing the interest of the live broadcast content. Moreover, the automatically and dynamically generated method can reduce manual intervention, making the comic style in the generated target video stream more natural and vivid; (2) By determining the updated target comic elements based on the selection information of the candidate comic elements, and generating an updated target video stream based on the updated target comic elements, a viewer interaction mechanism is provided to update the target video stream according to the viewer's selection, enhancing the viewer's sense of participation and interactivity during the live broadcast. Viewers can select and switch different candidate comic elements according to their own preferences, rather than being limited to a fixed comic style, comic special effects, or comic characters; (3) By determining the weight information of each candidate comic element based on the statistical quantity of the selection information of the candidate comic elements, and generating a comprehensive comic element as the updated target comic element based on the weight information, the generated comprehensive comic style not only includes each comic style but also reflects the voting preferences of the viewers. Thus, a richer and more personalized visual effect is provided. Compared with a single comic style, the comprehensive comic style can better meet the aesthetic needs of different viewers and present a comic style effect with mixed aesthetic characteristics; (4) By adding comic special effects matching the first information to the target video stream, the visual effect is further enhanced on the basis of the target video stream with the target comic style, creating the atmosphere of the live broadcast room and further enhancing the interest of the live broadcast content; (5) By fusing the first information with the original video frame, input data containing rich information can be generated, enabling the comic style conversion model to better understand and convert the original video frame and improving the conversion effect; (6) By adopting a generative adversarial network to implement the comic style conversion model, high-quality images can be output, ensuring the real-time performance and low latency rate of the output images; (7) By receiving the target video stream with the target comic style sent by the server side and displaying the target video stream with the target comic style in the live broadcast screen, the target video stream is generated by the server side based on converting the original video stream in the first information sent from the live broadcast end, and the target video stream has a target comic style matching the original video stream. Thus, a target video stream with a specific comic style matching the content of the original video stream is presented in the viewer-side live broadcast screen, making the visual effect of the live broadcast screen more rich and enhancing the interest of the live broadcast content.Moreover, the target video stream is automatically generated dynamically, which can reduce manual intervention and make the comic effects in the target video stream more natural and vivid. (8) By obtaining and sending the original video stream of the live broadcast screen to the server side, receiving the target video stream with the target comic style sent by the server side, and displaying the target video stream on the live broadcast screen, the target video stream can be generated by the server side based on the first information in the original video stream to convert the original video stream, and has the target comic style matching the original video stream, so that the target video stream with a specific comic style matching the content of the original video stream is presented in the live broadcast screen of the host side, making the visual effect of the live broadcast screen more rich and enhancing the interest of the live broadcast content. Moreover, the target video stream is automatically generated dynamically without the need for the host to manually adjust. This generation method can reduce manual intervention and make the comic effects in the target video stream more natural and vivid. It should be noted that different embodiments may produce different beneficial effects. In different embodiments, the beneficial effects that may be produced can be any one or several combinations of the above, or any other beneficial effects that may be obtained.
[0167] In some embodiments, the processor may be a combination of one or more of the following processors: central processing unit (CPU), application specific integrated circuit (ASIC), application specific instruction set processor (ASIP), graphics processing unit (GPU), physics processing unit (PPU), digital signal processor (DSP), field programmable gate array (FPGA), programmable logic device (PLD), programmable logic controller (PLC), reduced instruction set computer (RISC), microprocessor.
[0168] In some embodiments, the storage medium may include one or more combinations of the following: mass storage, removable storage, volatile read-write memory, read-only memory (ROM). Exemplary mass storage may include magnetic disks, optical disks, solid state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, compressed hard disks, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary random access memory may include dynamic random access memory (DRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), static random access memory (SRAM), thyristor random access memory (T-RAM), and zero capacitor memory (Z-RAM), etc. Exemplary read-only memory may include masked read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compressed hard disk read-only memory (CD-ROM), and digital versatile hard disk read-only memory, etc.
[0169] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are taught in this specification, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this specification.
Claims
1. A method for generating a live video, characterized in that, The method includes: Obtaining an original video stream; Identifying first information in the original video stream, where the first information includes at least one of human action information, human expression information, or scene information in the original video stream; Converting the original video stream based on the first information to generate a target video stream with a target comic style matching the original video stream.
2. The method according to claim 1, characterized in that It further includes: Obtaining selection information of candidate comic elements sent by the viewer side, where the candidate comic elements include at least one of a candidate comic style, candidate comic special effects, or candidate comic characters; Determining updated target comic elements based on the selection information of the candidate comic elements, and generating an updated target video stream based on the updated target comic elements.
3. The method according to claim 2, wherein When the selection information of the candidate comic elements includes the candidate comic style, the determining updated target comic elements based on the selection information of the candidate comic elements and generating an updated target video stream based on the updated target comic elements includes: Determining an updated target comic style based on the selection information of the candidate comic style; Converting the original video stream based on the updated target comic style to generate the updated target video stream, and the updated target video stream has the updated target comic style.
4. The method according to claim 2, wherein The determining updated target comic elements based on the selection information of the candidate comic elements includes: Determining the candidate comic element with the largest selection quantity as the updated target comic element based on the statistical quantity of the selection information of the candidate comic elements; or Determining weight information of each candidate comic element based on the statistical quantity of the selection information of the candidate comic elements, and generating a comprehensive comic element as the updated target comic element based on the weight information, where the comprehensive comic element integrates the characteristics of each candidate comic element.
5. The method according to claim 2, wherein When the selection information of the candidate comic elements includes the candidate comic special effects or the candidate comic characters, the determining updated target comic elements based on the selection information of the candidate comic elements and generating an updated target video stream based on the updated target comic elements includes: Determining target comic special effects or target comic characters based on the selection information of the candidate comic special effects or the candidate comic characters; Adjusting the target video stream with the target comic style based on the target comic special effects or the target comic characters to generate the updated target video stream, and the updated target video stream includes the target comic special effects or the target comic characters.
6. The method according to claim 5, wherein The adjusting the target video stream with the target comic style based on the target comic special effects or the target comic characters to generate the updated target video stream includes: Adding the target comic special effects to the target video stream with the target comic style; or Replacing the comic special effects in the target video stream with the target comic style with the target comic special effects; or Replacing the characters in the target video stream with the target comic style with the target comic characters.
7. The method according to claim 2, characterized in that, It further includes: Determine the candidate comic elements based on the preset information or the historical selection information of the viewer client, for the viewer client to select based on the interactive interface.
8. The method according to claim 1, characterized in that The human action information includes human key point information and / or human action categories; The identifying the first information in the original video stream includes: Identifying the human key point information in the original video stream; Determining the human action category in the original video stream based on the human key point information.
9. The method according to claim 1, wherein The human expression information includes at least one of human expression categories, facial key point information, and facial feature vectors; The identifying the first information in the original video stream includes: Identifying the facial key point information and / or the facial feature vector in the original video stream; Determining the human expression category in the original video stream based on the facial key point information and / or the facial feature vector.
10. The method according to claim 1, wherein The scene information includes at least one of human information, object information, and background information; The human information includes human position and human category, the object information includes object position and object category, and the background information includes background area and background category; Wherein, the human position includes human key point information; the object position includes the bounding box information of the object; the background area is determined based on the mask image of the background.
11. The method according to any one of claims 1 or 8 - 10, characterized in that, The original video stream is sent from the host client; The human action information includes the action information of the host; the human expression information includes the expression information of the host; the scene information includes the scene information of the live broadcast room where the host is located.
12. The method according to any one of claims 1 or 8 - 10, characterized in that, It further includes: Adding comic special effects matching the first information to the target video stream based on the first information.
13. The method according to claim 1, characterized in that The converting the original video stream based on the first information is implemented through a pre-trained comic style conversion model; The pre-trained comic style conversion model generates a target video stream with a target comic style matching the original video stream based on the first information and each original video frame in the original video stream, or based on the fusion feature map or fusion feature vector generated from the first information and each original video frame of the original video stream.
14. The method according to claim 13, wherein The comic style conversion model is trained based on a sample training set, and the sample training set includes video frame samples, first sample information extracted from the video frame samples, and comic style images matching the first sample information, and the first sample information includes at least one of action sample information, expression sample information, and scene sample information.
15. A live video display method, characterized in that, The method includes: Receiving a target video stream with a target comic style sent by the server, the target video stream being generated by the server converting the original video stream based on the first information in the original video stream sent from the host client, and the target video stream having the target comic style matching the original video stream; Displaying the target video stream with the target comic style on the live broadcast screen; Wherein, the first information includes at least one of human action information, human expression information, or scene information in the original video stream.
16. The method according to claim 15, wherein It further includes: Display candidate comic elements, where the candidate comic elements include at least one of a candidate comic style, candidate comic special effects, or candidate comic characters; Obtain selection information of the candidate comic elements input based on the interactive interface and send it to the server side; Receive and display the updated target video stream sent by the server side, where the updated target video stream is generated by the server side based on the selection information of the candidate comic elements.
17. The method according to claim 16, wherein, When the selection information of the candidate comic elements includes the candidate comic style, the updated target video stream has an updated target comic style determined based on the selection information of the candidate comic style; When the selection information of the candidate comic elements includes candidate comic special effects or candidate comic characters, the updated target video stream includes target comic special effects or target comic characters determined based on the selection information of the candidate comic special effects or the candidate comic characters.
18. The method according to claim 16, wherein, The updated target video stream is generated by the server side based on updated target comic elements, where the updated target comic elements are the candidate comic elements with the largest number of selections, or Composite comic elements, where the composite comic elements are generated based on the weight information of each candidate comic element.
19. A live video display method, characterized in that The method includes: Obtain and send the original video stream of the live broadcast screen to the server side; Receive the target video stream with the target comic style sent by the server side, where the target video stream is generated by the server side converting the original video stream based on the first information in the original video stream and has the target comic style matching the original video stream; Display the target video stream in the live broadcast screen; Wherein, the first information includes at least one of the human action information, human expression information, or scene information in the original video stream.
20. A live video generation device, characterized in that The device includes: An acquisition module for acquiring the original video stream; An identification module for identifying the first information in the original video stream, where the first information includes at least one of the human action information, human expression information, or scene information in the original video stream; A conversion module for converting the original video stream based on the first information to generate a target video stream with a target comic style matching the original video stream.
21. A live video display device, characterized in that, The device includes: A receiving module for receiving the target video stream with the target comic style sent by the server side, where the target video stream is generated by the server side converting the original video stream based on the first information in the original video stream sent from the live broadcast end, and the target video stream has the target comic style matching the original video stream; A display module for displaying the target video stream with the target comic style in the live broadcast screen; Wherein, the first information includes at least one of the human action information, human expression information, or scene information in the original video stream.
22. A live video display device, characterized in that, The device includes: An acquisition module for acquiring and sending the original video stream of the live broadcast screen to the server side; A receiving module, configured to receive a target video stream with a target comic style sent by the server side, where the target video stream is generated by the server side based on first information in the original video stream and has the target comic style matching the original video stream; A display module, configured to display the target video stream in the live broadcast screen; Wherein, the first information includes at least one of character action information, character expression information or scene information in the original video stream.
23. A live broadcast system, characterized in that, The system includes a server side, a host side and an audience side; The server side is configured to obtain an original video stream sent from the host side, identify the first information in the original video stream, convert the original video stream based on the first information, generate a target video stream with a target comic style matching the original video stream, and send it to the host side and the audience side; The host side is configured to obtain and send the original video stream to the server side, and obtain and display the target video stream with the target comic style sent by the server side; The audience side is configured to obtain and display the target video stream with the target comic style sent by the server side; Wherein, the first information includes at least one of character action information, character expression information or scene information in the original video stream.
24. A computer device, characterized in that, The computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the method according to any one of claims 1 to 19 can be implemented.
25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by the processor, the method according to any one of claims 1 to 19 can be implemented.
26. A computer program product, characterized in that, It includes a computer program, and when at least a part of the computer program is executed by the processor, the method according to any one of claims 1 to 19 can be implemented.