Method for generating information, method for displaying information, device, and storage medium
By generating virtual text and virtual objects and displaying video comments, the problems of low efficiency and poor performance of video comments in the prior art are solved, and a more intuitive and efficient user experience is achieved.
Patent Information
- Application Number
- PCT/CN2024/135785
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-30
- Filing Date
- 2024-11-29
- Publication Date
- 2025-06-05
AI Technical Summary
When presenting comments on product-type videos, existing video platforms have problems of low display efficiency and poor performance, and cannot effectively display valuable information in user comments to users intuitively.
By obtaining the description data corresponding to the video, a virtual text is generated using the language model. The virtual text comments on the video content by the virtual character, and a virtual object that is read aloud by voice, generating a second video to display the comment.
It improves the display efficiency and effect of video comments, enables users to understand the video content more intuitively, and enhances the user experience.
Smart Images

Figure CN2024135785_05062025_PF_FP_ABST
Abstract
Description
Information generation method, information display method, device and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202311633206.3 filed on November 30, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field
[0002] The embodiments of the present disclosure relate to an information generation method, an information display method, a device, and a storage medium. Background Art
[0003] Currently, after watching a product video on a video platform, users need to further check the comments about the product in order to further understand the product-related information.
[0004] For example, by responding to user operations, the terminal device can display the comment information corresponding to the video on the video page. However, the comment information on the video page has the problems of low display efficiency and poor display effect. Summary of the Invention
[0005] The embodiments of the present disclosure provide an information generation method, an information display method, a device, and a storage medium.
[0006] In a first aspect, an embodiment of the present disclosure provides an information generation method, comprising:
[0007] Obtain descriptive data corresponding to a first video, wherein the descriptive data represents historical comments on the first video; process the descriptive data through a language model to generate virtual text, wherein the content of the virtual text includes comments on target content in the first video based on the identity of a virtual character; create a virtual object that voice-reads the virtual text, and generate a second video based on the virtual object.
[0008] In a second aspect, an embodiment of the present disclosure provides an information display method, including:
[0009] Play a first video; in response to a trigger instruction for the first video, play a second video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0010] In a third aspect, an embodiment of the present disclosure provides an information generating device, including:
[0011] An acquisition module, configured to acquire description data corresponding to a first video, wherein the description data represents historical comments on the first video;
[0012] a processing module, processing the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on target content in the first video based on the identity of the virtual character;
[0013] A generation module is used to create a virtual object that voice-reads the virtual text, and generate a second video based on the virtual object.
[0014] In a fourth aspect, an embodiment of the present disclosure provides an information display device, including:
[0015] A first playback module, configured to play a first video;
[0016] The second playback module is used to play a second video in response to a trigger instruction for the first video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0017] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0018] The memory stores computer-executable instructions;
[0019] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the information generation method described in the first aspect and various possible designs of the first aspect, or executes the information display method described in the second aspect and various possible designs of the second aspect.
[0020] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the information generation method described in the first aspect and various possible designs of the first aspect is implemented, or the information display method described in the second aspect and various possible designs of the second aspect is implemented.
[0021] In the seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the information generation method described in the first aspect and various possible designs of the first aspect, or implements the information display method described in the second aspect and various possible designs of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0023] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;
[0024] FIG2 is a flow chart of the information display method according to an embodiment of the present disclosure;
[0025] FIG3 is a schematic diagram of a process of playing a second video provided by an embodiment of the present disclosure;
[0026] FIG4 is a second flow chart of the information display method provided by an embodiment of the present disclosure;
[0027] FIG5 is a flowchart of a specific implementation of step S204 in the embodiment shown in FIG2 ;
[0028] FIG6 is a schematic diagram of another process of playing a second video provided by an embodiment of the present disclosure;
[0029] FIG7 is a flowchart of another specific implementation of step S204 in the embodiment shown in FIG2 ;
[0030] FIG8 is a flowchart of a method for generating information according to an embodiment of the present disclosure;
[0031] FIG9 is a flowchart of a specific implementation of step S302 in the embodiment shown in FIG8 ;
[0032] FIG10 is a second flow chart of the information generation method provided in an embodiment of the present disclosure;
[0033] FIG11 is a flowchart of a specific implementation of step S402 in the embodiment shown in FIG10 ;
[0034] FIG12 is a flowchart of a specific implementation of step S404 in the embodiment shown in FIG10 ;
[0035] FIG13 is a structural block diagram of an information display device provided by an embodiment of the present disclosure;
[0036] FIG14 is a structural block diagram of an information generating device provided by an embodiment of the present disclosure;
[0037] FIG15 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure; and
[0038] FIG16 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0041] The following explains the application scenarios of the embodiments of the present disclosure:
[0042] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. The information generation method provided by the embodiment of the present disclosure can be applied to applications of video platforms with comment functions, such as short video applications, live broadcast applications, etc. More specifically, it can be applied to application scenarios in which video comments on a video are displayed to users. The execution subject of this embodiment can be a terminal device running the application of the video platform with comment functions, or a server deploying the server corresponding to the application, or other electronic devices that perform similar functions. Referring to FIG1 , taking a terminal device as an example, the terminal device is, for example, a smartphone, and a short video application client is running in the terminal device. By operating the terminal device, the user can play the target short video in the video playback page of the client. Later, when the user needs to view user comments on the video, the terminal device responds to the user operation, obtains the comment information from the server (application server) corresponding to the client, and displays the comment information on the comment page, thereby achieving the purpose of displaying the comment information.
[0043] The method of displaying comment information to users through the client running on the terminal device is usually to obtain the comment information from the server and display it in the form of text and images. However, due to the large number of comments, the complex content, and the large number of invalid comments, the simple listing and display of comment information cannot achieve efficient display of comment information. At the same time, the display effect is monotonous, making it difficult to intuitively display the valuable information in the user comments to the user on the terminal device. In other words, the comment information display is inefficient and poor.
[0044] The embodiments of the present disclosure provide an information display method to solve the above problems.
[0045] Referring to FIG2 , FIG2 is a flow chart of an information display method according to an embodiment of the present disclosure. The method of this embodiment can be applied in a terminal device, and the information display method includes:
[0046] Step S101: Play the first video.
[0047] Step S102: In response to a trigger instruction for the first video, a second video is played, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0048] For example, referring to the application scenario diagram shown in Figure 1, on the terminal device side, after the terminal device plays the first video on the video playback page through the client of the video application, the user operates the interactive control on the video playback page and inputs a trigger instruction for the first video to the terminal device. Specifically, for example, when the user clicks on the interactive component #1 on the video playback page, the terminal device receives the trigger instruction for the first video. Afterwards, the terminal device responds to the trigger instruction for the first video and plays the second video on the video playback page. The second video includes a virtual object that reads aloud virtual text, and the content of the virtual text is generated based on the historical comments of the first video.
[0049] FIG3 is a schematic diagram of a process for playing a second video provided by an embodiment of the present disclosure. As shown in FIG3 , a terminal device plays a first video in a window within a video playback page. Specifically, the video content played in the first video is, for example, "A video exploring XX restaurant." Subsequently, a user clicks an interactive component within the video playback page, generating a trigger instruction for the first video. As shown in the figure, the interactive component is named "Open Video Comments." The terminal device responds to the trigger instruction, retrieves a second video corresponding to the first video from a server, and plays the second video within a pop-up comment page. The second video includes a digital human (i.e., a virtual object) that reads aloud virtual text generated based on historical comments about the first video. The virtual text includes comments on the target content of the first video based on the virtual character's identity. Specifically, the virtual text read by the virtual object may read, for example, "I go to this restaurant often. It's relatively expensive, but the food is good." Through the above method, the video content of the second video is displayed. Of course, in other embodiments, the virtual object can also be other digital entities besides a digital human, such as digital animals, figurative objects, etc., and can be configured as needed without specific limitation.
[0050] Exemplarily, the virtual text is generated based on historical comments about the first video. In one possible implementation, the virtual text is the result of screening numerous historical comments. For example, after sorting them based on the number of likes, the virtual text is generated using the content of highly liked comments. In another possible implementation, the virtual text is generated based on a language model. That is, the historical comments about the first video are processed using a language model, and the generated original comments are used as virtual text. That is, the text content of the virtual text is different from any historical comments.
[0051] In this embodiment, by obtaining the second video corresponding to the first video and playing it, the purpose of explaining the video content of the first video in a virtual character identity in video form is achieved, anthropomorphic commentary on the video content is achieved, and the display efficiency and display effect of the video commentary are improved.
[0052] Further, referring to FIG4 , FIG4 is a second flow chart of the information display method provided by the embodiment of the present disclosure. Based on the embodiment shown in FIG2 , this embodiment further refines steps S101-S102, wherein a video playback page for playing a first video and a second video has a first window and a second window, wherein the first window and the second window have a layer relationship. The information display method includes:
[0053] Step S201: Obtain video stream data.
[0054] Step S202: Obtain a first video and a second video based on the video stream data.
[0055] Step S203: determining a layer relationship between the first window and the second window based on the video playback states corresponding to the first window and the second window.
[0056] Step S204: Play the first video in the first window in the video playback page.
[0057] Step S205: Play the second video in the second window in the video playback page.
[0058] Exemplarily, video stream data refers to data received by a terminal device from a service end (server) for playing a first video. More specifically, it refers to feed stream data. After receiving the video stream data, the terminal device decodes it to obtain the first video and the corresponding second video carried in the video stream data. The video stream data may include multiple data streams, and the first video and the second video may be provided in the same data stream or in different data streams. The specific implementation method is set as needed.
[0059] Next, the layer relationship between the first window playing the first video and the second window playing the second video is determined. A layer relationship refers to the hierarchical relationship between the layers in which the first and second windows reside, i.e., the mutual obstruction between the two windows. For example, if the first window is located in a higher-level layer and the second window is located in a lower-level layer, when the first and second windows interfere with each other, the higher-level first window obscures the lower-level second window, and vice versa. Video playback states include at least two states: "paused" and "playing." In one possible implementation, a window in the "playing" state is set to the higher level, and a window in the "paused" state is set to the lower level. That is, the window corresponding to the first or second video currently playing is placed in the upper layer. Specifically, for example, when the first video is playing and the second video is paused, the first window is displayed in the upper layer, while the second window is hidden in the lower layer. Conversely, when the first video is paused and the second video is playing, the second window is displayed in the upper layer, while the first window is hidden in the lower layer. At the same time, since the first video is the main content of the video playback page, in one possible implementation, the first window corresponding to the first video is the main window (larger window) of the video playback page, and the second window corresponding to the second video is the sub-window of the video playback page. When the layer where the second window is located is higher than the layer where the first window is located, the second window floats above the first window. At this time, the second window and the first window (the part not blocked by the second window) can be seen at the same time on the video playback page. When the layer where the second window is located is lower than the layer where the first window is located, the second window is located below the first window and is invisible.
[0060] Afterwards, the corresponding first video and second video are played in the first window and the second window, wherein the first video and the second video can be played simultaneously, stopped simultaneously, or played alternately. In subsequent embodiments, the alternating playback situation will be introduced in more detail.
[0061] Optionally, after step S204, the process may return to step S201, repeat the above steps, and update the layer relationship according to the video playback status, thereby achieving switching display between the first window and the second window. In this embodiment, by obtaining the video playback status of the first video and the second video, the layer relationship corresponding to the two is set, thereby achieving switching display between the two videos on the same video playback page.
[0062] In a possible implementation, as shown in FIG5 , a specific implementation of step S204 includes:
[0063] Step S2041: Obtaining time information corresponding to the first video through the video stream data, where the time information is used to represent a time period corresponding to at least one content segment in the first video;
[0064] Step S2042: Play the second video according to the time information.
[0065] Exemplarily, the video stream data obtained by the terminal device includes time information representing the time period corresponding to at least one content segment in the first video. A content segment is divided based on the video entry of the first video. For example, the first video contains three content segments: the first content segment introduces "product performance," the first content segment introduces "product price," and the first content segment introduces "product after-sales service." The start and end timestamps corresponding to each of these content segments constitute the time period corresponding to the content segment. Subsequently, based on this time information, the second video is played at intervals. Specifically, during the time period corresponding to content segment A in the first video, the second video is played; during the time period corresponding to content segment B in the first video, the second video is paused.
[0066] Specifically, in a possible implementation, the time information or video stream data further includes a content identifier representing the content value corresponding to the content segment. The target time period is determined by the content identifier, and the second video is played within the target time period.
[0067] FIG6 is a schematic diagram of another process for playing a second video provided by an embodiment of the present disclosure. Referring to FIG6 , after the terminal device obtains the time information Info corresponding to the first video through the video stream data, the terminal device divides the first video into content segment A, content segment B, and content segment C based on the time information Info, with corresponding time periods of (0, T1], (T1, T2], and (T2, T3), respectively. At the same time, the terminal device obtains the content identifier corresponding to each content segment. For example, the content identifier corresponding to content segment A and content segment C is L1, indicating high-value content, and the content identifier corresponding to content segment B is L0, indicating low-value content. Based on the above time information, during the time period of content segment A and content segment C corresponding to high-value content, the second video is paused, for example, as shown in the figure, the second window is hidden; and during the time period of content segment B corresponding to low-value content, the second window pops up and the second video is played in the second window. In this way, the purpose of playing the second video based on the content value interval is achieved, avoiding the impact of the playback of the second video on the information display of the first video and the impact of the playback of the first video on the information display of the second video; and improving the information display efficiency of the first and second videos.
[0068] In another possible implementation, the time information or video stream data also includes a feature identifier that represents the specific content corresponding to the content paragraph; the terminal device can determine the time period (of the content paragraph) corresponding to the target feature identifier as the target time period based on a preset feature mapping relationship, and then play the second video within the target time period, thereby achieving precise control of the playback timing of the second video, so that the playback timing of the second video can be accurately matched with the real-time playback content of the first video.
[0069] In another possible implementation, as shown in FIG7 , the specific implementation of step S204 includes:
[0070] Step S2043: Analyze the video content of the first video currently being played.
[0071] Step S2044: Play the second video according to the video content so that the voice speech in the second video and the voice speech in the first video are not played at the same time.
[0072] For example, in one possible implementation, the time information corresponding to the first video can be obtained by parsing the video content currently being played by the first video. The time information is used to represent the time period corresponding to at least one content segment in the first video. Then, the second video can be played according to the time information corresponding to the first video. Furthermore, more specifically, the voice speeches in the first video can be parsed based on the video content, and the time period of non-voice speeches can be determined as the target time period, while the time period of voice speeches can be determined as the non-target time period. Afterwards, the playback of the second video can be controlled based on the target time period, thereby achieving the purpose of playing the voice speeches in the second video at different times from the voice speeches in the first video.
[0073] Among them, after determining the target time period based on the video content, the way of controlling the playback of the first video and the second video based on the target time period is similar to the implementation method in the embodiment shown in Figure 6, and will not be repeated here.
[0074] With respect to the above-mentioned method executed on the terminal device side, the embodiment of the present disclosure also provides an information generation method to solve the above-mentioned problem. The method can be applied to a server outside the terminal device to generate a second video played on the terminal device. Of course, it can also be applied to the terminal device itself, that is, the second video in the above-mentioned embodiment is generated by the terminal device.
[0075] Referring to FIG8 , FIG8 is a flow chart of the information generation method provided in an embodiment of the present disclosure. The method of this embodiment can be applied in a server, and the information generation method includes:
[0076] Step S301: Obtain description data corresponding to a first video, where the description data represents historical comments on the first video;
[0077] For example, referring to the application scenario diagram shown in FIG1 and the process of the terminal device playing the second video in the embodiment shown in FIG2 , the second video played by the terminal device can be generated based on the information generation method provided in this embodiment. The execution subject of the method in this embodiment is a server, for example, a data generation server for generating a video. The data generation server can be connected to an application server running a server end, and send the generated second video to the application server, which then sends it to the terminal device for playback; the data generation server can also directly send the generated second video to the terminal device for playback, or the data generation server and the application server can be the same server.
[0078] Taking the case where the data generation server and the application server are the same server (referred to as the server in this embodiment) as an example, specifically, in order to generate a second video corresponding to a first video, the server first obtains description data corresponding to the first video, where the description data represents historical comments on the first video. For example, historical comments refer to submitted user comments on the video content, and the description data includes at least one of the following: text, images, videos, and voice. The description data can be stored locally on the server or in other external storage media. The specific method for obtaining the description data and the specific content of the description data are not further described here.
[0079] Step S302: Processing the description data using a language model to generate virtual text, where the content of the virtual text includes a comment on the target content in the first video based on the identity of the virtual character;
[0080] Afterwards, the server calls the language model to process the description data to extract the valid information in the above description data. Afterwards, based on the extracted valid information, processing is performed, such as summarizing and generalizing, to generate virtual text. The content of the virtual text includes comments on the target content in the first video based on the virtual character identity. Specifically, the language model includes, for example, a large language model (LLM) of various implementation methods, which completes the generation of the above virtual text based on the content generation (Generated Content) capability of the language model. Among them, the virtual text is a meaningful text segment composed of multiple characters. The content it represents is the comment on the target content in the first video by the virtual character identity. It is a text information that describes the content in the first person (i.e., the virtual character person). Its purpose is to simulate a user's speech. Afterwards, combined with the subsequent step of generating a second video containing a virtual object, the simulation of "real person speech and comments" is achieved.
[0081] Furthermore, in a possible implementation, as shown in FIG9 , a specific implementation of step S302 includes:
[0082] Step S3021: Obtain corresponding target prompt words based on the identification information of the first video, where the target prompt words are used to describe evaluation requirements for at least one evaluation dimension based on natural language;
[0083] Step S3022: Process the target prompt words and description data using the language model to generate virtual text.
[0084] Exemplarily, first, the server obtains the identification information of the first video, wherein the identification information may be the video identifier of the first video, more specifically, such as the video name, video number, etc.; the identification information may also be the category video information of the first video, such as the content category of the first video, for example, the identification information info_1 indicates that the first video is a product promotion video, and the identification information info_2 indicates that the first video is a film and television work video, wherein the identification information may be pre-bound to the first video, and the specific implementation method of the identification information may be set based on needs, and will not be described one by one here. Afterwards, based on the identification information, it is mapped to the corresponding prompt word template, wherein the prompt word template is a template that represents the format, content and other information of the prompt word, and is used to generate the target prompt word. The prompt word template contains a specific content format for the identification information, so that the target prompt word generated based on the prompt word template can describe the evaluation requirements for at least one evaluation dimension, for example, the content category represented by the identification information info_3 is "food exploration video". The prompt word template M_3 obtained based on the identification information info_3 contains prompt word parameters representing the "price dimension," "taste dimension," and "queue event dimension," as well as fixed text for evaluating the restaurant using these prompt word parameters (evaluation dimension). Subsequently, based on this prompt word template, the corresponding target prompt word can be generated. The mapping relationship between the identification information and the prompt word template can be pre-set and is not limited. Once the prompt word template is obtained, the specific content of the prompt word template and the specific implementation method for generating the corresponding prompt word based on the prompt word template are not further described here.
[0085] Of course, the target prompt word may also include other information, which can be flexibly set manually by the user and will not be described in detail here.
[0086] Afterwards, the target prompt words and description data are input into the language model, and the output of the language model is guided by the target prompt words to obtain the required content generation (AIGC) data, that is, virtual text.
[0087] Furthermore, in a possible implementation, the target prompt word includes role information, and the role information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character; accordingly, the target prompt word and the description data are processed using a language model to generate virtual text. The specific implementation method of step S3022 includes: using a language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0088] For example, in order to further generate more realistic and personalized virtual text so that the virtual object that subsequently speaks using the virtual text is more realistic, in this embodiment, role information is provided in the target prompt word, and the role information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character.
[0089] Among them, for the case where the role information is used to describe the identity characteristics of the virtual character, specifically, for example, the text content carrying the role information in the target prompt word is: "You are a chef", or "You are a young man from Province X". Accordingly, based on the target prompt word containing the above role information, it is possible to guide the language model to comment on the target content of the first video from different "identity characteristics" dimensions, thereby generating different results (virtual texts). More specifically, for example, when the text content of the role information is: "You are a chef", the corresponding generated virtual text includes the following content: "I am a chef, and my evaluation of the taste of this restaurant is...". Correspondingly, when the text content of the role information is: "You are a young man from Province X", the corresponding generated virtual text includes the following content: "I am a young man from Province X, and my evaluation of the taste of this restaurant is..."
[0090] In another possible implementation, for the case where the character information is used to describe the preference characteristics of the virtual character, specifically, for example, the text content carrying the character information in the target prompt word is: "You like to go out to eat after 8 pm", or "You like lighter food". Accordingly, based on the target prompt word containing the above character information, it is possible to guide the language model to comment on the target content of the first video from different "preference characteristics" dimensions, thereby generating different results (virtual texts). More specifically, for example, when the text content of the character information is: "You like to go out to eat after 8 pm", the corresponding generated virtual text includes the following content: "I like to go out to eat after 8 pm, and the queue situation at this restaurant is...". Correspondingly, when the text content of the character information is: "You like lighter food", the corresponding generated virtual text includes the following content: "I like lighter food, and my evaluation of the taste of this restaurant is..."
[0091] In this embodiment, by setting role information in the target prompt word, the target prompt word can guide the language model to comment on the target content of the first video from different role dimensions, which is equivalent to pre-setting the conditions for evaluating the target content, thereby making the generated virtual text more accurate and sufficient, thereby improving the information display efficiency of the second video.
[0092] Step S303: creating a virtual object for voice reading of the virtual text, and generating a second video based on the virtual object.
[0093] For example, after obtaining the generated virtual text, a virtual object, such as a digital human, is created to read the virtual text aloud, thereby simulating a "real person's speech and comments." On the one hand, by creating a model of the virtual object and driving the virtual object based on the virtual text obtained in the above steps, the virtual object's movements match the content of the virtual text. Specifically, for example, the virtual object is a pre-created digital human, and the digital human is driven by the virtual text so that the digital human's mouth movements match the content of the virtual text. This will not be described in detail here. Of course, in other possible implementations, other movements can also be generated by driving the virtual object, for example, driving the digital human's arms and limbs to produce movements that match the content of the virtual text. This will not be described in detail here.
[0094] On the other hand, the virtual text is converted into corresponding voice data, and the voice data is played synchronously while driving the above-mentioned virtual object action, thereby achieving the effect of driving the virtual object to read the virtual text aloud. In one possible implementation, for example, after creating the virtual object, the virtual text is converted into corresponding voice data through a preset voice generation model, and the voice data is input into the initial model of the above-mentioned virtual object to drive the virtual object to produce the corresponding voice reading action.
[0095] Of course, in another possible implementation, virtual text can be directly input into the model of the virtual object to directly realize the voice reading of the virtual object. This depends on the specific implementation of the virtual object model and can be set as needed. Afterwards, by recording the virtual object with the virtual text, the corresponding second video can be obtained. This will not be repeated here.
[0096] Furthermore, optionally, after creating a virtual object that reads the virtual text aloud, the background of the virtual object can be further configured to control the background content of the generated second video. Specifically, the method of this embodiment further includes:
[0097] Acquire semantic features corresponding to the virtual text; obtain corresponding video background material according to the semantic features; generate a second video based on the virtual object, including: generating a second video based on the virtual object and the video background material.
[0098] Exemplarily, the virtual text is further processed to extract semantic features corresponding to the virtual text. Semantic features can be features used to characterize the subject content of the virtual text, implemented in the form of a matrix or text. Subsequently, based on the semantic features, material generation is performed to obtain video background material that matches the content of the first video. The video background material can be a generated image obtained locally or externally after matching based on the semantic features, or it can be an image generated using AIGC technology based on the semantic features. There is no limitation here. In this embodiment, the semantic features of the virtual text are used to match the corresponding video background material for the virtual object, and then a second video is generated based on the video background material and the virtual object, thereby improving the visual performance of the second video.
[0099] Afterwards, after receiving the request for the second video sent by the terminal device, the server sends the second video to the terminal device, thereby achieving the playback of the second video on the terminal device.
[0100] In this embodiment, descriptive data corresponding to a first video is obtained, representing historical comments on the first video. The descriptive data is processed using a language model to generate virtual text, which includes comments on target content in the first video based on the identity of a virtual character. A virtual object is created that reads the virtual text aloud, and based on the virtual object, a second video is generated and played on the terminal device. By converting the historical comments on the first video into a second video based on a language model and playing it back, the historical comments on the first video, after being understood and organized by the language model, are presented in the form of a video, making the video content more concise and effective, and improving the efficiency and effectiveness of the display of comment information.
[0101] Referring to FIG. 10 , FIG. 10 is a second flow chart of the information generation method provided in an embodiment of the present disclosure. Based on the embodiment shown in FIG. 8 , this embodiment further refines step S303. The information generation method includes:
[0102] Step S401: Obtain description data corresponding to a first video, where the description data represents historical comments on the first video.
[0103] Step S402: Process the description data through a language model to generate virtual text, where the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0104] Step S403: Create a virtual object.
[0105] For example, the virtual object can be a 3D model of a digital human. After pre-training, this 3D model can receive voice data and control the digital human to make at least one of matching mouth movements, facial expressions, and body movements. The specific training and use process of the model is not detailed here.
[0106] Furthermore, the three-dimensional model corresponding to the virtual object may be fixed or dynamically changing. In one possible implementation, as shown in FIG11 , the specific implementation of step S403 includes:
[0107] Step S4031: Generate feature information of a virtual object based on the video content of the first video, where the feature information is used to characterize the appearance features of the virtual object.
[0108] Step S4032: Create a virtual object based on the feature information.
[0109] For example, generating feature information of a virtual object based on the video content of the first video can be accomplished by parsing the first video to obtain a feature matrix representing the video content of the first video, processing the feature matrix and classifying it to obtain a corresponding content identifier. Subsequently, based on the content identifier, the feature information is mapped to the corresponding appearance features representing the virtual object. Specifically, for example, when the video content of the first video is classified as "football match video," the appearance features (feature information) of the corresponding virtual object are "a person wearing a football uniform." The aforementioned appearance features can also be represented by specific feature identifiers, which will not be described in detail.
[0110] In the steps of this embodiment, by analyzing the video content of the first video, matching feature information is generated, and then the appearance features of the virtual object are dynamically determined, so that the appearance of the virtual object in the second video can match the video content of the first video, thereby improving the consistency between the two.
[0111] Step S404: Generate corresponding voice data according to the virtual text.
[0112] Exemplarily, a preset speech generation model can be used to convert virtual text into corresponding speech data, wherein the speech generation model has at least one input parameter, namely, virtual text. In one possible implementation, the virtual text is used as an input parameter, such as a speech generation model, to obtain corresponding speech data. This case will not be described in detail. In other possible implementations, the speech generation model also includes at least one control parameter for controlling the sound characteristics of the generated speech data. As shown in FIG12 , exemplarily, the specific implementation of step S404 includes:
[0113] Step S4041: Acquire feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object.
[0114] Step S4042: Based on the feature information, the virtual text is processed to generate voice data having sound features that match the appearance features.
[0115] Exemplarily, in combination with the embodiment shown in Figure 11, after generating the feature information of the virtual object, the virtual text can be further processed based on the feature information of the virtual object to generate voice data with sound features that match the appearance features. For example, the appearance features of the virtual object generated based on the video content of the first video are "young men"; then, the feature information is used as a control parameter, and the voice generation model is input to process the virtual text to generate voice data with sound features that match the appearance features, and the voice data has the voice features of a young man.
[0116] In this embodiment, by determining feature information based on the video content of the first video, voice data that matches it is generated, so that the appearance features of the virtual object can match the voice features when it is read aloud, thereby improving the authenticity and consistency of the virtual object in the second video and improving the video performance effect.
[0117] Step S405: Based on the voice data, drive the virtual object to perform corresponding actions and / or expressions, and save them as a second video.
[0118] In this embodiment, the specific implementation of steps S401, S402, and S405 has been described in detail in the embodiment shown in FIG7 and will not be repeated here.
[0119] Corresponding to the information generation method of the above embodiment, FIG13 is a structural block diagram of the information display device provided by the embodiment of the present disclosure. For the sake of convenience, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG13, the information display device 5 includes:
[0120] A first playback module 51, configured to play a first video;
[0121] The second playback module 52 is used to play the second video in response to the trigger instruction for the first video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0122] According to one or more embodiments of the present disclosure, the first playback module 51 is specifically used to: obtain video stream data; obtain a first video based on the video stream data, and play the first video in a first window within a video playback page; when playing the second video, the second playback module 52 is specifically used to: obtain a second video based on the video stream data, and play the second video in a second window within the video playback page; wherein the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback status corresponding to the first window and the second window.
[0123] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module 52 is specifically used to: obtain time information corresponding to the first video through video stream data, where the time information is used to represent the time period corresponding to at least one content paragraph in the first video; and play the second video according to the time information.
[0124] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module 52 is specifically used to: analyze the video content currently played by the first video; play the second video according to the video content so that the voice speech in the second video and the voice speech in the first video are not played at the same time.
[0125] The first playback module 51 and the second playback module 52 are connected. The information display device 5 provided in this embodiment can implement the technical solution of the corresponding embodiment of the above-mentioned information display method, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.
[0126] Corresponding to the information generation method of the above embodiment, FIG14 is a structural block diagram of the information generation device provided by the embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG14, the information generation device 6 includes:
[0127] An acquisition module 61 is configured to acquire description data corresponding to the first video, where the description data represents historical comments on the first video;
[0128] The processing module 62 processes the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character;
[0129] The generating module 63 is used to create a virtual object for voice reading of the virtual text, and generate a second video based on the virtual object.
[0130] According to one or more embodiments of the present disclosure, the processing module 62 is specifically used to: obtain corresponding target prompt words based on the recognition information of the first video, and the target prompt words are used to describe the evaluation requirements for at least one evaluation dimension based on natural language; use the language model to process the target prompt words and description data to generate virtual text.
[0131] According to one or more embodiments of the present disclosure, the target prompt words include role information, and the role information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character; when the processing module 62 uses the language model to process the target prompt words and description data to generate virtual text, it is specifically used to: use the language model to process the target prompt words and description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0132] According to one or more embodiments of the present disclosure, the generation module 63 is specifically used to: create a virtual object; generate corresponding voice data based on the virtual text; drive the virtual object to perform corresponding actions and / or expressions based on the voice data, and save it as a second video.
[0133] According to one or more embodiments of the present disclosure, when creating a virtual object, the generation module 63 is specifically used to: generate feature information of the virtual object based on the video content of the first video, where the feature information is used to characterize the appearance features of the virtual object; and create the virtual object based on the feature information.
[0134] According to one or more embodiments of the present disclosure, when the generation module 63 generates corresponding voice data based on virtual text, it is specifically used to: obtain feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object; based on the feature information, process the virtual text to generate voice data with sound features that match the appearance features.
[0135] According to one or more embodiments of the present disclosure, the processing module 62 is also used to: obtain semantic features corresponding to the virtual text; obtain corresponding video background materials based on the semantic features; when the generation module 63 generates a second video based on the virtual object, it is specifically used to: generate a second video based on the virtual object and the video background material.
[0136] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, and voice.
[0137] The acquisition module 61, processing module 62 and generation module 63 are connected in sequence. The information generation device 6 provided in this embodiment can implement the technical solution of the above-mentioned information generation method embodiment, and its implementation principle and technical effect are similar, which will not be repeated in this embodiment.
[0138] FIG15 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. As shown in FIG15 , the electronic device 7 includes:
[0139] A processor 71, and a memory 72 communicatively connected to the processor 71;
[0140] Memory 72 stores computer-executable instructions;
[0141] The processor 71 executes the computer-executable instructions stored in the memory 72 to implement the information display method in the embodiments shown in Figures 2 to 7, or to implement the information generation method in the embodiments shown in Figures 8 to 12.
[0142] Optionally, the processor 71 and the memory 72 are connected via a bus 73 .
[0143] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the embodiments corresponding to Figures 2 to 12, and no further details will be given here.
[0144] An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the information display method provided in any of the embodiments corresponding to Figures 2 to 7 of the present disclosure, or to implement the information generation method provided in any of the embodiments corresponding to Figures 8 to 12 of the present disclosure.
[0145] An embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the information display method provided in any one of the embodiments corresponding to Figures 2 to 7 of the present disclosure, or is used to implement the information generation method provided in any one of the embodiments corresponding to Figures 8 to 12 of the present disclosure.
[0146] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.
[0147] Referring to FIG16 , a schematic diagram of the structure of an electronic device 900 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG16 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.
[0148] As shown in FIG16 , the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0149] Typically, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG16 shows an electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.
[0150] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0151] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.
[0152] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0153] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.
[0154] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).
[0155] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0156] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."
[0157] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0158] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0159] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an information generation method, comprising:
[0160] Obtain description data corresponding to a first video, wherein the description data represents historical comments on the first video; process the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on the target content in the first video based on the identity of a virtual character; create a virtual object that reads the virtual text aloud, and generate a second video based on the virtual object.
[0161] According to one or more embodiments of the present disclosure, the processing of the description data through a language model to generate virtual text includes: obtaining corresponding target prompt words based on the identification information of the first video, the target prompt words being used to describe the evaluation requirements for at least one evaluation dimension based on natural language; and using the language model to process the target prompt words and the description data to generate virtual text.
[0162] According to one or more embodiments of the present disclosure, the target prompt word includes character information, and the character information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character; using the language model to process the target prompt word and the description data to generate virtual text includes: using the language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0163] According to one or more embodiments of the present disclosure, a virtual object that voice-reads the virtual text is created, and a second video is generated based on the virtual object, including: creating a virtual object; generating corresponding voice data based on the virtual text; and driving the virtual object to perform corresponding actions and / or expressions based on the voice data, and saving the results as the second video.
[0164] According to one or more embodiments of the present disclosure, the creating of the virtual object includes: generating feature information of the virtual object based on the video content of the first video, wherein the feature information is used to characterize the appearance features of the virtual object; and creating the virtual object based on the feature information.
[0165] According to one or more embodiments of the present disclosure, generating corresponding voice data based on the virtual text includes: obtaining feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object; and processing the virtual text based on the feature information to generate voice data having sound features that match the appearance features.
[0166] According to one or more embodiments of the present disclosure, the method further includes: obtaining semantic features corresponding to the virtual text; obtaining corresponding video background materials based on the semantic features; generating a second video based on the virtual object includes: generating a second video based on the virtual object and the video background materials.
[0167] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, and voice.
[0168] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an information display method, comprising:
[0169] Play a first video; in response to a trigger instruction for the first video, play a second video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0170] According to one or more embodiments of the present disclosure, the playing of the first video includes: obtaining video stream data; obtaining the first video based on the video stream data, and playing the first video in a first window within the video playback page; the playing of the second video includes: obtaining the second video based on the video stream data, and playing the second video in a second window within the video playback page; wherein the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback status corresponding to the first window and the second window.
[0171] According to one or more embodiments of the present disclosure, playing the second video includes: obtaining time information corresponding to the first video through video stream data, the time information is used to represent the time period corresponding to at least one content segment in the first video; and playing the second video according to the time information.
[0172] According to one or more embodiments of the present disclosure, playing the second video includes: parsing the video content currently played by the first video; playing the second video according to the video content, so that the voice speech in the second video and the voice speech in the first video are not played at the same time.
[0173] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an information generating device, comprising:
[0174] An acquisition module, configured to acquire description data corresponding to a first video, wherein the description data represents historical comments on the first video;
[0175] a processing module, processing the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on target content in the first video based on the identity of the virtual character;
[0176] A generation module is used to create a virtual object that voice-reads the virtual text, and generate a second video based on the virtual object.
[0177] According to one or more embodiments of the present disclosure, the processing module is specifically used to: obtain corresponding target prompt words based on the identification information of the first video, and the target prompt words are used to describe the evaluation requirements for at least one evaluation dimension based on natural language; use the language model to process the target prompt words and the description data to generate virtual text.
[0178] According to one or more embodiments of the present disclosure, the target prompt word includes character information, and the character information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character; when the processing module uses the language model to process the target prompt word and the description data to generate virtual text, it is specifically used to: use the language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0179] According to one or more embodiments of the present disclosure, the generation module is specifically used to: create a virtual object; generate corresponding voice data based on the virtual text; drive the virtual object to perform corresponding actions and / or expressions based on the voice data, and save it as the second video.
[0180] According to one or more embodiments of the present disclosure, when creating a virtual object, the generation module is specifically used to: generate feature information of the virtual object based on the video content of the first video, wherein the feature information is used to characterize the appearance features of the virtual object; and create the virtual object based on the feature information.
[0181] According to one or more embodiments of the present disclosure, when the generation module generates corresponding voice data based on the virtual text, it is specifically used to: obtain feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object; based on the feature information, process the virtual text to generate voice data with sound features that match the appearance features.
[0182] According to one or more embodiments of the present disclosure, the processing module is further used to: obtain semantic features corresponding to the virtual text; obtain corresponding video background materials based on the semantic features; when the generation module generates a second video based on the virtual object, it is specifically used to: generate a second video based on the virtual object and the video background material.
[0183] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, and voice.
[0184] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided an information display device, comprising:
[0185] A first playback module, configured to play a first video;
[0186] The second playback module is used to play a second video in response to a trigger instruction for the first video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
[0187] According to one or more embodiments of the present disclosure, the first playback module is specifically used to: obtain video stream data; obtain the first video based on the video stream data, and play the first video in a first window within the video playback page; when playing the second video, the second playback module is specifically used to: obtain the second video based on the video stream data, and play the second video in a second window within the video playback page; wherein, the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback status corresponding to the first window and the second window.
[0188] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module is specifically used to: obtain time information corresponding to the first video through video stream data, and the time information is used to represent the time period corresponding to at least one content paragraph in the first video; play the second video according to the time information.
[0189] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module is specifically used to: parse the video content currently played by the first video; and play the second video according to the video content so that the voice speech in the second video and the voice speech in the first video are not played at the same time.
[0190] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;
[0191] The memory stores computer-executable instructions;
[0192] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the information generation method described in the first aspect and various possible designs of the first aspect, or executes the information display method described in the second aspect and various possible designs of the second aspect.
[0193] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the information generation method described in the first aspect and various possible designs of the first aspect is implemented, or the information display method described in the second aspect and various possible designs of the second aspect is implemented.
[0194] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the information generation method described in the first aspect and various possible designs of the first aspect, or implements the information display method described in the second aspect and various possible designs of the second aspect.
[0195] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0196] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0197] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating information, comprising: Obtaining description data corresponding to a first video, wherein the description data represents historical comments on the first video; Processing the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character; A virtual object is created to voice read the virtual text, and a second video is generated based on the virtual object.
2. The method according to claim 1, wherein: The step of processing the description data through a language model to generate virtual text includes: Obtaining a corresponding target prompt word according to the identification information of the first video, where the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language; The target prompt word and the description data are processed by using the language model to generate the virtual text.
3. The method according to claim 2, wherein: The target prompt word includes role information, and the role information is used to describe the identity characteristics of the virtual character and / or the preference characteristics of the virtual character; Processing the target prompt word and the description data using the language model to generate the virtual text includes: The target prompt word and the description data are processed using the language model to generate virtual text corresponding to the identity feature and / or preference feature.
4. The method according to any one of claims 1 to 3, wherein: Creating a virtual object that voice-reads the virtual text, and generating the second video based on the virtual object, includes: Create virtual objects; According to the virtual text, generating corresponding voice data; Based on the voice data, the virtual object is driven to perform corresponding actions and / or expressions, and saved as the second video.
5. The method according to claim 4, wherein: The creating of the virtual object comprises: generating feature information of a virtual object according to the video content of the first video, wherein the feature information is used to characterize the appearance features of the virtual object; The virtual object is created according to the feature information.
6. The method according to claim 4 or 5, wherein: The step of generating corresponding voice data according to the virtual text includes: Acquire feature information of the virtual object, where the feature information is used to characterize appearance features of the virtual object; Based on the feature information, the virtual text is processed to generate voice data having a sound feature matching the appearance feature.
7. The method according to any one of claims 1 to 6, further comprising: Acquire semantic features corresponding to the virtual text; According to the semantic features, corresponding video background material is obtained; The step of generating a second video based on the virtual object comprises: The second video is generated based on the virtual object and the video background material.
8. The method according to any one of claims 1 to 7, wherein: The description data includes at least one of the following: Text, pictures, videos, voice.
9. An information display method, comprising: Play the first video; In response to a trigger instruction for the first video, a second video is played, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on target content in the first video based on the identity of the virtual character.
10. The method according to claim 9, wherein: The playing of the first video includes: Get video stream data; Based on the video stream data, the first video is obtained, and the first video is played in a first window in a video playback page; The playing of the second video comprises: Based on the video stream data, obtaining the second video, and playing the second video in a second window in the video playback page; The first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback status corresponding to the first window and the second window.
11. The method according to claim 9, wherein: Playing the second video includes: Obtaining time information corresponding to the first video through the video stream data, where the time information is used to represent a time period corresponding to at least one content segment in the first video; Play the second video according to the time information.
12. The method according to claim 9, wherein: Playing the second video includes: Analyze the video content currently being played by the first video; The second video is played according to the video content so that the voice speech in the second video is not played simultaneously with the voice speech in the first video.
13. An information generating device, comprising: An acquisition module is configured to acquire description data corresponding to a first video, wherein the description data represents historical comments on the first video; A processing module, which processes the description data through a language model to generate virtual text, wherein the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character; The generating module is configured to create a virtual object for voice reading the virtual text, and generate a second video based on the virtual object.
14. An information display device, comprising: A first playback module, configured to play a first video; The second playback module is configured to play a second video in response to a trigger instruction for the first video, wherein the second video includes a virtual object that voice-reads virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character.
15. An electronic device comprising a processor and a memory, wherein: The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the information generating method according to any one of claims 1 to 8, or executes the information display method according to any one of claims 9 to 12.
16. A computer-readable storage medium storing computer-executable instructions, wherein: When the processor executes the computer-executable instruction, the information generation method as described in any one of claims 1 to 8 is implemented, or the information display method as described in any one of claims 9 to 12 is implemented.
17. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, it implements the information generation method according to any one of claims 1 to 8, or implements the information display method according to any one of claims 9 to 12.
Citation Information
Patent Citations
Commenting method and device, and electronic equipment
CN108322832A
Video processing method based on commodity object, electronic equipment and storage medium
CN115690275A
Virtual robot interaction method and device and storage medium
CN116756285A
Virtual influencers for narration of spectated video games
US20210322888A1
Cited By
Method, equipment and device for displaying cultural relics based on XR
CN120355874A