Information generation method, information display method, equipment and storage medium
By generating and displaying virtual comment videos based on virtual character identities, the problem of low efficiency and poor performance of comment information display within the video page is solved, and a more streamlined and effective comment display is achieved.
Patent Information
- Application Number
- CN202311633206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
In the prior art, the comment information in the video page has problems such as low display efficiency and poor display effect.
By obtaining the description data corresponding to the first video, processing these data using a language model, generating virtual text. The virtual text content comments on the target content of the first video based on the identity of the virtual character, and creates a virtual object that reads the virtual text aloud, and finally generates a second video.
By converting historical comments into video formats, the streamlining and effective display of comment information is achieved, and the efficiency and effectiveness of comment information is improved.
Smart Images

Figure CN120075517A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of Internet technologies, and in particular, to an information generation method, an information display method, a device, and a storage medium. Background Art
[0002] Currently, after a user watches a product-related video on a video platform, there is a need to further view the comments on the product, so as to further understand the relevant information of the product.
[0003] In the prior art, by responding to a user operation, a terminal device can display the comment information corresponding to a video on the video page. However, in the solutions of the prior art, the comment information in the video page has problems of low display efficiency and poor display effect. Summary of the Invention
[0004] Embodiments of the present disclosure provide an information generation method, an information display method, a device, and a storage medium to overcome the problems of low display efficiency and poor display effect of comment information.
[0005] In a first aspect, embodiments of the present disclosure provide an information generation method, including:
[0006] Obtaining description data corresponding to a first video, where the description data represents historical comments of the first video; processing the description data through a language model to generate virtual text, where the content of the virtual text includes comments on target content in the first video based on a virtual character identity; creating a virtual object that reads the virtual text aloud, and generating a second video based on the virtual object.
[0007] In a second aspect, embodiments of the present disclosure provide an information display method, including:
[0008] Playing a first video; in response to a trigger instruction for the first video, playing a second video, where the second video includes a virtual object that reads virtual text aloud, and the content of the virtual text includes comments on target content in the first video based on a virtual character identity.
[0009] In a third aspect, embodiments of the present disclosure provide an information generation device, including:
[0010] An obtaining module, configured to obtain description data corresponding to a first video, where the description data represents historical comments of the first video;
[0011] A processing module, configured to process the description data through a language model to generate virtual text, where the content of the virtual text includes comments on target content in the first video based on a virtual character identity;
[0012] A generation module, configured to create a virtual object for reading the virtual text aloud and generate a second video based on the virtual object.
[0013] In a fourth aspect, an embodiment of the present disclosure provides an information display device, including:
[0014] A first playback module, configured to play a first video;
[0015] A second playback module, configured to play a second video in response to a trigger instruction for the first video, where the second video includes a virtual object that reads the virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
[0016] In a fifth aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;
[0017] The memory stores computer-executable instructions;
[0018] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the information generation method described in the first aspect above and various possible designs of the first aspect, or executes the information display method described in the second aspect above and various possible designs of the second aspect.
[0019] In a sixth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the information generation method described in the first aspect above and various possible designs of the first aspect is implemented, or the information display method described in the second aspect above and various possible designs of the second aspect is implemented.
[0020] In a seventh aspect, an embodiment of the present disclosure provides a computer program product, including a computer program. When the computer program is executed by a processor, the information generation method described in the first aspect above and various possible designs of the first aspect is implemented, or the information display method described in the second aspect above and various possible designs of the second aspect is implemented.
[0021] The information generation method, information display method, device, and storage medium provided in this embodiment obtain description data corresponding to a first video, where the description data represents historical comments on the first video; process the description data through a language model to generate virtual text, and the content of the virtual text includes comments on target content in the first video based on the identity of a virtual character; create a virtual object that reads the virtual text aloud, generate a second video based on the virtual object, and play the second video on a terminal device. By converting the historical comments of the first video into a second video and playing it based on the language model, the historical comments of the first video are presented in the form of a video after being understood and organized by the language model, so that the video content is more concise and effective, and the display efficiency and display effect of comment information are improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following briefly introduces the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0023] Figure 1 FIG. is a schematic diagram of an application scenario provided for an embodiment of the present disclosure;
[0024] Figure 2 FIG. is a flowchart of the information display method provided for an embodiment of the present disclosure Figure 1 ;
[0025] Figure 3 FIG. is a schematic diagram of the process of playing a second video provided for an embodiment of the present disclosure;
[0026] Figure 4 FIG. is a flowchart of the information display method provided for an embodiment of the present disclosure Figure 2 ;
[0027] Figure 5 For Figure 2 FIG. is a flowchart of a specific implementation manner of step S204 in the shown embodiment;
[0028] Figure 6 FIG. is another schematic diagram of the process of playing a second video provided for an embodiment of the present disclosure;
[0029] Figure 7 For Figure 2 FIG. is a flowchart of another specific implementation manner of step S204 in the shown embodiment;
[0030] Figure 8Flow schematic of the information generation method provided by the embodiments of the present disclosure Figure 1 ;
[0031] Figure 9 For Figure 8 Flowchart of the specific implementation manner of step S302 in the illustrated embodiment;
[0032] Figure 10 Flow schematic of the information generation method provided by the embodiments of the present disclosure Figure 2 ;
[0033] Figure 11 For Figure 10 Flowchart of the specific implementation manner of step S402 in the illustrated embodiment;
[0034] Figure 12 For Figure 10 Flowchart of the specific implementation manner of step S404 in the illustrated embodiment;
[0035] Figure 13 Block diagram of the information display device provided by the embodiments of the present disclosure;
[0036] Figure 14 Block diagram of the information generation device provided by the embodiments of the present disclosure;
[0037] Figure 15 Structural schematic diagram of an electronic device provided by the embodiments of the present disclosure;
[0038] Figure 16 Hardware structural schematic diagram of the electronic device provided by the embodiments of the present disclosure. Specific embodiments
[0039] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0040] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0041] The application scenarios of the embodiments of the present disclosure are explained below:
[0042] Figure 1 FIG. is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. The information generation method provided by the embodiment of the present disclosure can be applied to an application program of a video platform with a comment function, such as a short video application, a live broadcast application, etc. More specifically, it can be applied to an application scenario of displaying video comments for a video to a user. The execution subject of this embodiment can be a terminal device running the application program of the above-mentioned video platform with a comment function, or a server deploying the server corresponding to the above application program, or other electronic devices with similar functions. Refer to Figure 1 As shown in, taking the terminal device as an example, the terminal device is, for example, a smart phone. A short video application client is running in the terminal device. By operating the terminal device, the user can play a target short video on the video playback page of the client. After that, when the user needs to view the user comments on the video, the terminal device responds to the user operation, obtains comment information from the server (application server) corresponding to the client, and displays the comment information on the comment page to achieve the purpose of displaying the comment information.
[0043] In the prior art, the method of displaying comment information to a user through a client running in a terminal device is usually to obtain comment information from a server and display it in the form of text and pictures. However, due to the large number and complex content of comment information, including a large number of invalid comments, the simple listing and display method of comment information in the prior art cannot achieve the efficient display of comment information, and at the same time, the display effect is single, and it is difficult to intuitively display the valuable information in the user comments to the user on the terminal device side, that is, there are problems of low display efficiency and poor display effect of comment information.
[0044] The embodiments of the present disclosure provide an information display method to solve the above problems.
[0045] Refer to Figure 2 , Figure 2 is a flow chart of the information display method provided by an embodiment of the present disclosure Figure 1 . The method of this embodiment can be applied to a terminal device. The information display method includes:
[0046] Step S101: Play the first video.
[0047] Step S102: In response to a trigger instruction for the first video, play a second video, where the second video includes a virtual object that reads virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
[0048] Exemplarily, refer to Figure 1The schematic diagram of the application scenario shown. On the terminal device side, after the terminal device plays the first video in the video playback page through the client of the video application, the user operates on the interactive control in the video playback page and inputs a trigger instruction for the first video to the terminal device. Specifically, for example, when the user clicks on the interactive component #1 in the video playback page, the terminal device receives the trigger instruction for the first video. After that, the terminal device responds to the trigger instruction for the first video and plays the second video in the video playback page. Among them, the second video includes a virtual object that reads virtual text aloud, and the content of the virtual text is generated based on the historical comments of the first video.
[0049] Figure 3 The schematic diagram of the process of playing the second video provided by an embodiment of the present disclosure is as Figure 3 shown. The terminal device plays the first video in the window in the video playback page. Specifically, the video content played in the first video is, for example, "a video of exploring XX restaurant". After that, the user clicks on the interactive component in the video playback page to generate a trigger instruction for the first video. As shown in the figure, the name of the interactive component is "Open Video Comments". The terminal device responds to the trigger instruction, obtains the second video corresponding to the first video from the server, and plays the second video in the popped-up comment page. The second video includes a digital human (i.e., a virtual object), and the digital human reads aloud the virtual text generated based on the historical comments of the first video. The content of the virtual text includes comments on the target content in the first video based on the virtual character identity. Specifically, the content of the virtual text read by the virtual object is, for example, "I often go to this restaurant. The price is relatively expensive, but the taste is also good". Through the above method, the display of the video content of the second video is realized. Of course, in other embodiments, the virtual object can also be other digital bodies other than digital humans, such as digital animals, image objects, etc., which can be set according to needs and are not specifically limited.
[0050] Among them, exemplarily, the virtual text is generated based on the historical comments of the first video. In one possible implementation, the virtual text is the screening result of numerous historical comments. For example, after sorting based on the "number of likes", the high-like comment content is used to generate the virtual text. In another possible implementation, the virtual text is generated based on a language model, that is, the historical comments of the first video are processed by the language model, and the generated original comments are used as the virtual text, that is, the text content of the virtual text is not the same as any historical comment.
[0051] In this embodiment, by obtaining the second video corresponding to the first video and playing it, the purpose of explaining the video content of the first video in the form of a virtual character identity is achieved, the anthropomorphic comment on the video content is realized, and the display efficiency and display effect of the video comment are improved.
[0052] Further, referring to Figure 4 , Figure 4 is a flowchart of the information display method provided by an embodiment of the present disclosure. Figure 2 . On the basis of the embodiment shown in Figure 2 , steps S101 - S102 are further refined. Among them, in the video playback page for playing the first video and the second video, there are a first window and a second window, and there is a layer relationship between the first window and the second window. The information display method includes:
[0053] Step S201: Obtain video stream data.
[0054] Step S202: Based on the video stream data, obtain the first video and the second video.
[0055] Step S203: Determine the layer relationship between the first window and the second window based on the video playback states corresponding to the first window and the second window.
[0056] Step S204: Play the first video in the first window within the video playback page.
[0057] Step S205: Play the second video in the second window within the video playback page.
[0058] Exemplarily, the video stream data refers to the data received by the terminal device from the server side for realizing the playback of the first video. More specifically, for example, feed stream data. After receiving the video stream data, the terminal device decodes it to obtain the first video and the corresponding second video carried in the video stream data. Among them, the video stream data can include multiple data streams. The first video and the second video can be set in the same data stream or in different data streams respectively, and the specific implementation method is set according to needs.
[0059] After that, determine the layer relationship between the first window for playing the first video and the second window for playing the second video. Here, the layer relationship refers to the upper and lower hierarchical relationship of the layers where the first window and the second window are located, that is, the mutual occlusion relationship between the two. For example, if the first window is on the high-level layer and the second window is on the low-level layer, when the first window and the second window interfere with each other, the high-level first window occludes the low-level second window, and vice versa. The video playback state includes at least two states: "paused" and "playing". In a possible implementation, the window in the "playing" state is set to the high level, and the window in the "paused" state is set to the low level. That is, for the first video and the second video, the corresponding window of the one that is playing is on the upper layer. Specifically, for example, when the first video is playing and the second video is paused, the first window is displayed on the upper layer, and the second window is hidden on the lower layer; when the first video is paused and the second video is playing, the second window is displayed on the upper layer, and the first window is hidden on the lower layer. At the same time, since the first video is the main playback content of the video playback page, in a possible implementation, the first window corresponding to the first video is the main window (larger window) of the video playback page, and the second window corresponding to the second video is the secondary window of the video playback page. When the layer where the second window is located is higher than the layer where the first window is located, the second window floats above the first window. At this time, both the second window and the first window (the part not occluded by the second window) can be seen within the video playback page. When the layer where the second window is located is lower than the layer where the first window is located, the second window is located below the first window and is not visible.
[0060] After that, play the corresponding first video and second video in the first window and the second window. Here, the first video and the second video can be played simultaneously, stopped simultaneously, or played alternately. In subsequent embodiments, the case of alternate play will be introduced in more detail.
[0061] Optionally, after step S204, it is possible to return to step S201, repeat the above steps, and update the layer relationship according to the video playback state, so as to realize the switching display between the first window and the second window. In the steps of this embodiment, by obtaining the video playback states of the first video and the second video, the layer relationship corresponding to the two is set, so as to realize the mutual switching display of the two videos within the same video playback page.
[0062] In a possible implementation, as Figure 5 shown, the specific implementation manner of step S204 includes:
[0063] Step S2041: Obtain the time information corresponding to the first video from the video stream data, where the time information is used to represent the time period corresponding to at least one content segment in the first video.
[0064] Step S2042: Play the second video according to the time information.
[0065] Exemplarily, in the video stream data obtained by the terminal device, there is time information used to represent the time period corresponding to at least one content segment in the first video. Here, the content segment is a segment divided based on the entry of the first video. For example, the first video contains 3 content segments. The first content segment introduces "product performance", the first content segment introduces "product price", and the first content segment introduces "product after-sales". The start and end timestamps corresponding to each of the above content segments constitute the time period corresponding to the content segment. Then, based on this time information, the second video is played at intervals. Specifically, when the first video is played to the time period corresponding to content segment A, the second video is played; when the first video is played to the time period corresponding to content segment B, the second video is paused.
[0066] Specifically, in a possible implementation manner, the time information or the video stream data further includes a content identifier representing the content value corresponding to the content segment. The target time period is determined through the content identifier, and the second video is played within the target time period.
[0067] Figure 6 Another schematic diagram of the process of playing the second video provided by the embodiments of the present disclosure is referred to Figure 6 As shown, after the terminal device obtains the time information Info corresponding to the first video from the video stream data, based on this time information Info, the first video is divided into content segment A, content segment B, and content segment C, and the corresponding time periods are (0, T1], (T1, T2], and (T2, T3] respectively. At the same time, the content identifiers corresponding to each content segment are obtained. For example, the content identifiers corresponding to content segment A and content segment C are L1, indicating high-value content, and the content identifier corresponding to content segment B is L0, indicating low-value content. According to the above time information, within the time periods of content segment A and content segment C corresponding to high-value content, the playback of the second video is paused. For example, as shown in the figure, the second window is hidden; while within the time period of content segment B corresponding to low-value content, the second window is popped up and the second video is played within the second window. Thus, the purpose of playing the second video at intervals based on the content value is achieved, avoiding the influence of the playback of the second video on the information display of the first video and the influence of the playback of the first video on the information display of the second video; improving the information display efficiency of the first video and the second video.
[0068] In another possible implementation, the time information or video stream data further includes a feature identifier representing the specific content corresponding to the content segment; the terminal device may, based on a preset feature mapping relationship, determine the time period corresponding to the target feature identifier (of the content segment) as the target time period, and then play the second video within the target time period, thereby achieving precise control over the playback timing of the second video and enabling the playback timing of the second video to precisely match the real-time playback content of the first video.
[0069] In another possible implementation, as Figure 7 shown, the specific implementation of step S204 includes:
[0070] Step S2043: Analyze the video content currently being played in the first video.
[0071] Step S2044: Play the second video according to the video content so that the voice speech in the second video and the voice speech in the first video are not played simultaneously.
[0072] Exemplarily, in one possible implementation, the time information corresponding to the first video can be obtained by analyzing the video content currently being played in the first video. The time information is used to represent the time periods corresponding to at least one content segment in the first video. Further, more specifically, the voice speech in the first video can be parsed according to the video content, and the time periods of non-voice speech are determined as the target time periods, while the time periods of voice speech are determined as non-target time periods. Then, the playback of the second video is controlled based on the target time periods, thereby achieving the purpose that the voice speech in the second video and the voice speech in the first video are not played simultaneously.
[0073] Among them, after determining the target time period based on the video content, the manner of controlling the playback of the first video and the second video based on the target time period is Figure 6 similar to the implementation manner in the embodiments shown and will not be elaborated here one by one.
[0074] Regarding the method executed on the terminal device side above, correspondingly, the embodiments of the present disclosure further provide an information generation method to solve the above problems. This method can be applied to a server external to the terminal device to generate the second video played on the terminal device. Of course, it can also be applied to the terminal device itself, that is, the terminal device generates the second video in the above embodiments.
[0075] Refer to Figure 8 , Figure 8 which is the flowchart of the information generation method provided by the embodiments of the present disclosure Figure 1 . The method of this embodiment can be applied in a server. The information generation method includes:
[0076] Step S301: Obtain the description data corresponding to the first video, where the description data represents the historical comments of the first video;
[0077] Exemplarily, refer to Figure 1 the schematic diagram of the application scenario shown, and Figure 2 the process of the terminal device playing the second video in the shown embodiment. The second video played by the terminal device can be generated based on the information generation method provided in this embodiment. The execution subject of the method in this embodiment is the server, for example, a data generation server for generating videos. This data generation server can be connected to the application server running the server side, and send the generated second video to the application server, and then the application server sends it to the terminal device for playing; it can also be directly sent by the data generation server to the terminal device for playing. Or, the data generation server and the application server are the same server.
[0078] Taking the case where the data generation server and the application server are the same server (abbreviated as the server in this embodiment) as an example, specifically, in order to generate the second video corresponding to the first video, the server first obtains the description data corresponding to the first video, where the description data represents the historical comments of the first video. Exemplarily, the historical comments refer to the user comments that have been submitted for the video content. The description data includes at least one of the following: text, picture, video, voice. The description data can be stored locally on the server or other external storage media. The specific way to obtain the description data and the specific content of the description data are not elaborated here.
[0079] Step S302: Process the description data through a language model to generate virtual text, where the content of the virtual text includes comments on the target content in the first video based on the virtual character identity;
[0080] After that, the server calls the language model to process the description data, extracts the valid information in the above description data, and then processes the extracted valid information, such as summarizing and generalizing, to generate virtual text. The content of the virtual text includes comments on the target content in the first video based on the virtual character's identity. Specifically, the language model includes, for example, large language models (LLMs) with various implementation methods, and the virtual text is generated based on the content generation ability of the language model. Among them, the virtual text is a meaningful text segment composed of multiple words. The content it represents is a comment on the target content in the first video from the perspective of a virtual character, which is a text information described in the first person (i.e., the virtual character's person), aiming to simulate a user's speech. Then, combined with the subsequent step of generating the second video containing virtual objects, it realizes the simulation of "real person speech comments".
[0081] Further, in a possible implementation, as Figure 9 shown, the specific implementation of step S302 includes:
[0082] Step S3021: Obtain the corresponding target prompt word according to the recognition information of the first video, where the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language;
[0083] Step S3022: Use the language model to process the target prompt word and the description data to generate virtual text.
[0084] Exemplarily, first, the server obtains the identification information of the first video. Here, the identification information may be the video identifier of the first video. More specifically, for example, the video name, video number, etc.; the identification information may also be the category video information of the first video, such as the content category of the first video. For example, the identification information info_1 indicates that the first video is a product promotion video, and the identification information info_2 indicates that the first video is a film and television work video. Here, the identification information may be pre-bound to the first video, and the specific implementation manner of the identification information may be set based on requirements, and will not be elaborated one by one here. After that, based on this identification information, it is mapped to the corresponding prompt template. Here, the prompt template is a template that represents information such as the format and content of the prompt, and is used to generate the target prompt. The prompt template contains a specific content format for the identification information, so that the target prompt generated based on this prompt template can describe the evaluation requirements for at least one evaluation dimension. For example, the content category represented by the identification information info_3 is "food tasting and restaurant visiting video". Then, the prompt template M_3 obtained according to this identification information info_3 contains prompt parameters representing "price dimension", "taste dimension", and "queueing event dimension", as well as fixed text for combining the above prompt parameters (evaluation dimensions) to evaluate a "restaurant". After that, based on the above prompt template, the corresponding target prompt can be generated. The mapping relationship between the identification information and the prompt template can be preset in advance without limitation; after obtaining the prompt template, the specific content of the prompt template and the specific implementation manner of generating the corresponding prompt based on the prompt template will not be elaborated here.
[0085] Of course, the target prompt may also include other information, which can be flexibly set manually by the user and will not be elaborated here.
[0086] After that, the obtained target prompt and description data are input into the language model, and the output of the language model is guided by the target prompt, and the required content generation (AIGC) data, that is, virtual text, can be obtained.
[0087] Further, in a possible implementation manner, the target prompt includes role information, and the role information is used to describe the identity characteristics of the virtual role and / or the preference characteristics of the virtual role; correspondingly, the specific implementation manner of using the language model to process the target prompt and description data to generate virtual text in step S3022 includes: using the language model to process the target prompt and description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0088] Exemplarily, in order to further generate more realistic and personalized virtual text to make the virtual object that uses the virtual text to speak more realistic, in this embodiment, role information is set in the target prompt. The role information is used to describe the identity characteristics of the virtual role and / or the preference characteristics of the virtual role.
[0089] Among them, for the case where the role information is used to describe the identity characteristics of the virtual role, specifically, for example, the text content carrying the role information in the target prompt is: "You are a chef", or "You are a young person from Province X". Correspondingly, based on the target prompt containing the above role information, it is possible to guide the language model to comment on the target content of the first video from different "identity characteristic" dimensions, thereby generating different results (virtual text). More specifically, for example, when the text content of the role information is: "You are a chef", the virtual text generated correspondingly contains the following content: "I am a chef, and my evaluation of the taste of this restaurant is..." Correspondingly, when the text content of the role information is: "You are a young person from Province X", the virtual text generated correspondingly contains the following content: "I am a young person from Province X, and my evaluation of the taste of this restaurant is..."
[0090] In another possible implementation, for the case where the role information is used to describe the preference characteristics of the virtual role, specifically, for example, the text content carrying the role information in the target prompt is: "You like to go out for dinner after 8 pm", or "You like lighter food". Correspondingly, based on the target prompt containing the above role information, it is possible to guide the language model to comment on the target content of the first video from different "preference characteristic" dimensions, thereby generating different results (virtual text). More specifically, for example, when the text content of the role information is: "You like to go out for dinner after 8 pm", the virtual text generated correspondingly contains the following content: "I like to go out for dinner after 8 pm, and the queue situation in this restaurant is..." Correspondingly, when the text content of the role information is: "You like lighter food", the virtual text generated correspondingly contains the following content: "I like lighter food, and my evaluation of the taste of this restaurant is..."
[0091] In this embodiment, by setting role information in the target prompt, the target prompt can guide the language model to comment on the target content of the first video from different role dimensions, which is equivalent to pre-setting the conditions for evaluating the target content, so that the generated virtual text is more accurate and sufficient, thereby improving the information display efficiency of the second video.
[0092] Step S303: Create a virtual object for voice reading of the virtual text, and generate a second video based on the virtual object.
[0093] Exemplarily, further, after obtaining the generated virtual text, create a virtual object, such as a digital human, to perform language reading on the virtual text, so as to simulate the "real person's speech comment". On the one hand, by creating a model of the virtual object and driving the virtual object based on the virtual text obtained in the above steps, the actions of the virtual object are matched with the content of the virtual text. Specifically, for example, the virtual object is a pre-created digital human, and the digital human is driven by the virtual text, so that the mouth movements of the digital human are matched with the content of the virtual text. The specific implementation method is the prior art and will not be elaborated here. Of course, in other possible implementation manners, other actions of the virtual object can also be driven. For example, the arms and limbs of the digital human can be driven to generate movements that match the content of the virtual text, which will not be elaborated one by one here.
[0094] On the other hand, convert the virtual text into corresponding voice data, and while driving the actions of the above virtual object, synchronously play the voice data, so as to achieve the effect of driving the virtual object to voice read the virtual text. In one possible implementation manner, exemplarily, after creating the virtual object, convert the virtual text into corresponding voice data through a preset voice generation model, and then input the voice data into the initial model of the above virtual object to drive the virtual object to generate corresponding voice reading actions.
[0095] Of course, in another possible implementation manner, the virtual text can also be directly input into the model of the virtual object to directly achieve the voice reading of the virtual object, which depends on the specific implementation manner of the model of the virtual object and can be set according to needs here. After that, by recording the virtual object that voices the virtual text, the corresponding second video can be obtained, which will not be elaborated here.
[0096] Further, optionally, after creating the virtual object for voice reading the virtual text, the background of the virtual object can be further configured to control the background content of the generated second video. Specifically, the method of this embodiment further includes:
[0097] Obtain the semantic features corresponding to the virtual text; according to the semantic features, obtain the corresponding video background materials; generate a second video based on the virtual object, including: generating a second video based on the virtual object and the video background materials.
[0098] Exemplarily, the virtual text is further processed to extract semantic features corresponding to the virtual text. The semantic features can be features used to characterize the theme content of the virtual text and are implemented in the form of a matrix or text. Then, based on the semantic features, material generation is performed to obtain video background materials that match the content of the first video. The video background materials can be generated images obtained locally or externally after matching based on the semantic features, or can be images generated using AIGC technology based on the semantic features, and there is no limitation here. In this embodiment, through the semantic features of the virtual text, corresponding video background materials are matched for the virtual object, and then based on the video background materials and the virtual object, the second video is generated, improving the visual performance effect of the second video.
[0099] After that, after receiving the request for the second video sent by the terminal device, the server sends the second video to the terminal device, thereby realizing the playback of the second video on the terminal device.
[0100] In this embodiment, by obtaining the description data corresponding to the first video, the description data represents the historical comments of the first video; through the language model, the description data is processed to generate virtual text, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity; a virtual object for voice-reading the virtual text is created, and based on the virtual object, the second video is generated and the second video is played on the terminal device side. By converting the historical comments of the first video into the second video and playing it through the language model, the historical comments of the first video are understood and sorted by the language model and then presented in the form of a video, so that the video content is more concise and effective, improving the display efficiency and display effect of the comment information.
[0101] Reference Figure 10 , Figure 10 is the flowchart of the information generation method provided by the embodiment of the present disclosure Figure 2 In this embodiment, on the basis of the embodiment shown in Figure 8 further refine step S303. The information generation method includes:
[0102] Step S401: Obtain the description data corresponding to the first video, and the description data represents the historical comments of the first video.
[0103] Step S402: Through the language model, process the description data to generate virtual text, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
[0104] Step S403: Create a virtual object.
[0105] Exemplarily, the virtual object may be a three-dimensional model of a digital human. After being pre-trained, the three-dimensional model can receive voice data and control the digital human to perform at least one of matching mouth movements, facial expressions, and body movements. The specific training and use processes of the model are prior arts known to those skilled in the art and will not be elaborated herein.
[0106] Further, among them, the three-dimensional model corresponding to the virtual object may be fixed or dynamically changing. In one possible implementation, as Figure 11 shown, the specific implementation manner of step S403 includes:
[0107] Step S4031: Generate feature information of the virtual object according to the video content of the first video, where the feature information is used to characterize the appearance features of the virtual object.
[0108] Step S4032: Create a virtual object according to the feature information.
[0109] Among them, exemplarily, generating the feature information of the virtual object according to the video content of the first video may be to parse the first video to obtain a feature matrix representing the video content of the first video, process the feature matrix, classify to obtain the corresponding content identifier, and then map based on the content identifier to the corresponding appearance features representing the virtual object. Specifically, for example, when the video content of the first video is classified as a "football game video", the corresponding appearance features (feature information) of the virtual object are "people wearing football uniforms". The above appearance features can also be represented by specific feature identifiers and will not be elaborated herein.
[0110] In the steps of this embodiment, by parsing the video content of the first video to generate matching feature information, and then dynamically determining the appearance features of the virtual object, the appearance of the virtual object in the second video can be matched with the video content of the first video, improving the consistency between the two.
[0111] Step S404: Generate corresponding voice data according to the virtual text.
[0112] Exemplarily, using a preset voice generation model, the virtual text can be converted into corresponding voice data. Among them, the voice generation model has at least one input parameter, that is, the virtual text. In one possible implementation, taking the virtual text as the input parameter, for example, the voice generation model, the corresponding voice data can be obtained, and this situation will not be elaborated herein. In other possible implementation manners, the voice generation model further includes at least one control parameter for controlling the sound features of the generated voice data. As Figure 12 shown, exemplarily, the specific implementation manner of step S404 includes:
[0113] Step S4041: Obtain the feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object.
[0114] Step S4042: Process the virtual text based on the feature information to generate voice data with voice features matching the appearance features.
[0115] Exemplarily, in combination with Figure 11 the embodiments shown, after generating the feature information of the virtual object, the virtual text can be further processed based on the feature information of the virtual object to generate voice data with voice features matching the appearance features. For example, based on the video content of the first video, the appearance feature of the virtual object is "young man"; then, using this feature information as a control parameter, input it into the voice generation model to process the virtual text and generate voice data with voice features matching the appearance features, and this voice data has the voice features of a young man.
[0116] In this embodiment, by generating the feature information determined according to the video content of the first video to generate the matching voice data, the appearance features of the virtual object can be matched with the voice features during its voice reading, thereby improving the authenticity and consistency of the virtual object in the second video and enhancing the video presentation effect.
[0117] Step S405: Drive the virtual object to perform corresponding actions and / or expressions based on the voice data and save it as the second video.
[0118] In this embodiment, the specific implementation manners of steps S401, S402, and S405 are Figure 7 already introduced in detail in the embodiments shown and will not be elaborated here.
[0119] Corresponding to the information generation method in the above embodiments, Figure 13 this is the structural block diagram of the information display device provided by the embodiments of the present disclosure. For the sake of clarity, only the parts related to the embodiments of the present disclosure are shown.
[0120] Referring to Figure 13 , the information display device 5 includes:
[0121] A first playback module 51 for playing the first video;
[0122] A second playback module 52 for playing the second video in response to a trigger instruction for the first video, where the second video includes a virtual object that reads the virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
[0123] According to one or more embodiments of the present disclosure, the first playback module 51 is specifically configured to: obtain video stream data; based on the video stream data, obtain a first video, and play the first video in a first window within a video playback page; when the second playback module 52 plays a second video, it is specifically configured to: based on the video stream data, obtain a second video, and play the second video in a second window within the video playback page; wherein, the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback states corresponding to the first window and the second window.
[0124] According to one or more embodiments of the present disclosure, when the second playback module 52 plays a second video, it is specifically configured to: through the video stream data, obtain the time information corresponding to the first video, and the time information is used to represent the time period corresponding to at least one content segment in the first video; according to the time information, play the second video.
[0125] According to one or more embodiments of the present disclosure, when the second playback module 52 plays a second video, it is specifically configured to: parse the video content currently played by the first video; according to the video content, play the second video so that the voice speeches in the second video and the first video are not played simultaneously.
[0126] Wherein, the first playback module 51 and the second playback module 52 are connected. The information display device 5 provided in this embodiment can execute the technical solutions of the corresponding embodiments of the above information display method, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.
[0127] Corresponding to the information generation method in the above embodiment, Figure 14 is a structural block diagram of the information generation device provided in the embodiments of the present disclosure. For the sake of convenience of description, only the parts related to the embodiments of the present disclosure are shown.
[0128] Referring to Figure 14 , the information generation device 6 includes:
[0129] An acquisition module 61, configured to acquire description data corresponding to a first video, and the description data represents the historical comments of the first video;
[0130] A processing module 62, which processes the description data through a language model to generate virtual text, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity;
[0131] A generation module 63, configured to create a virtual object for voice reading the virtual text, and generate a second video based on the virtual object.
[0132] According to one or more embodiments of the present disclosure, the processing module 62 is specifically configured to: obtain a corresponding target prompt word according to the recognition information of the first video, where the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language; process the target prompt word and the description data by using a language model to generate virtual text.
[0133] According to one or more embodiments of the present disclosure, the target prompt word includes role information, where the role information is used to describe the identity characteristics of the virtual role and / or the preference characteristics of the virtual role; when the processing module 62 processes the target prompt word and the description data by using a language model to generate virtual text, it is specifically configured to: process the target prompt word and the description data by using a language model to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0134] According to one or more embodiments of the present disclosure, the generation module 63 is specifically configured to: create a virtual object; generate corresponding voice data according to the virtual text; drive the virtual object to perform corresponding actions and / or expressions based on the voice data, and save it as a second video.
[0135] According to one or more embodiments of the present disclosure, when the generation module 63 creates a virtual object, it is specifically configured to: generate feature information of the virtual object according to the video content of the first video, where the feature information is used to characterize the appearance characteristics of the virtual object; create a virtual object according to the feature information.
[0136] According to one or more embodiments of the present disclosure, when the generation module 63 generates corresponding voice data according to the virtual text, it is specifically configured to: obtain the feature information of the virtual object, where the feature information is used to characterize the appearance characteristics of the virtual object; process the virtual text based on the feature information to generate voice data with sound characteristics matching the appearance characteristics.
[0137] According to one or more embodiments of the present disclosure, the processing module 62 is further configured to: obtain the semantic feature corresponding to the virtual text; obtain the corresponding video background material according to the semantic feature; when the generation module 63 generates a second video based on the virtual object, it is specifically configured to: generate a second video based on the virtual object and the video background material.
[0138] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, voice.
[0139] Among them, the acquisition module 61, the processing module 62, and the generation module 63 are connected in sequence. The information generation device 6 provided in this embodiment can execute the technical solutions of the information generation method embodiment described above, and its implementation principle and technical effects are similar, and will not be elaborated here in this embodiment.
[0140] Figure 15A schematic structural diagram of an electronic device provided by an embodiment of the present disclosure is shown as follows Figure 15 As shown, the electronic device 7 includes:
[0141] A processor 71 and a memory 72 communicatively connected to the processor 71;
[0142] The memory 72 stores computer-executable instructions;
[0143] The processor 71 executes the computer-executable instructions stored in the memory 72 to implement the information display method in the embodiment shown as Figures 2 - 7 shown, or to implement the information generation method in the embodiment shown as Figures 8 - 12 shown.
[0144] Optionally, the processor 71 and the memory 72 are connected through a bus 73.
[0145] For relevant descriptions, reference can be made to the relevant descriptions and effects corresponding to the steps in the Figures 2 - 12 corresponding embodiments, and details are not elaborated here.
[0146] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the information display method provided by any one of the embodiments corresponding to the present disclosure Figures 2 - 7 when executed by a processor, or to implement the information generation method provided by any one of the embodiments corresponding to the present disclosure Figures 8 - 12 when executed by a processor.
[0147] An embodiment of the present disclosure provides a computer program product including a computer program, which implements the information display method provided by any one of the embodiments corresponding to the present disclosure Figures 2 - 7 when executed by a processor, or to implement the information generation method provided by any one of the embodiments corresponding to the present disclosure Figures 8 - 12 when executed by a processor.
[0148] To implement the above embodiments, an embodiment of the present disclosure further provides an electronic device.
[0149] Refer to Figure 16, which shows a schematic structural diagram of an electronic device 900 suitable for implementing the embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. Among them, the terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (PADs), portable media players (PMPs), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 16 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0150] As Figure 16 shown, the electronic device 900 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0151] Generally, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 16 the electronic device 900 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices can be implemented or had.
[0152] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product that includes a computer program carried on a computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0153] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0154] The above-mentioned computer-readable medium can be included in the above-mentioned electronic device; or it can exist separately without being assembled into the electronic device.
[0155] The above-mentioned computer-readable medium carries one or more programs, and when the above-mentioned one or more programs are executed by the electronic device, the electronic device is caused to perform the methods shown in the above embodiments.
[0156] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any kind of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0157] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0158] The units described in the embodiments of the present disclosure may be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit may also be described as "the unit for acquiring at least two Internet protocol addresses".
[0159] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, by way of non-limitation, exemplary types of hardware logic components that may be used include: Field Programmable Gate Arrays (FPGA), Application Specific Integrated Circuits (ASIC), Application Specific Standard Products (ASSP), Systems on Chip (SOC), Complex Programmable Logic Devices (CPLD), and so on.
[0160] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0161] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an information generation method, including:
[0162] Obtaining description data corresponding to a first video, where the description data characterizes historical comments of the first video; processing the description data through a language model to generate virtual text, the content of the virtual text including comments on target content in the first video based on the identity of a virtual character; creating a virtual object that reads the virtual text aloud, and generating a second video based on the virtual object
[0163] According to one or more embodiments of the present disclosure, the processing the description data through a language model to generate virtual text includes: obtaining a corresponding target prompt word according to the identification information of the first video, where the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language; using the language model to process the target prompt word and the description data to generate virtual text.
[0164] According to one or more embodiments of the present disclosure, the target prompt word includes character information, where the character information is used to describe the identity characteristics and / or preference characteristics of the virtual character; using the language model to process the target prompt word and the description data to generate virtual text includes: using the language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0165] According to one or more embodiments of the present disclosure, creating a virtual object that reads the virtual text aloud and generating a second video based on the virtual object includes: creating a virtual object; generating corresponding voice data according to the virtual text; driving the virtual object to perform corresponding actions and / or expressions based on the voice data, and saving it as the second video.
[0166] According to one or more embodiments of the present disclosure, creating the virtual object includes: generating feature information of the virtual object according to the video content of the first video, where the feature information is used to characterize the appearance features of the virtual object; and creating the virtual object according to the feature information.
[0167] According to one or more embodiments of the present disclosure, generating corresponding voice data according to the virtual text includes: obtaining the feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object; and processing the virtual text based on the feature information to generate voice data having voice features matching the appearance features.
[0168] According to one or more embodiments of the present disclosure, the method further includes: obtaining semantic features corresponding to the virtual text; obtaining corresponding video background materials according to the semantic features; and generating a second video based on the virtual object, including: generating a second video based on the virtual object and the video background materials.
[0169] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, and voice.
[0170] In a second aspect, according to one or more embodiments of the present disclosure, an information display method is provided, including:
[0171] Playing a first video; in response to a trigger instruction for the first video, playing a second video, where the second video includes a virtual object that reads the virtual text aloud, and the content of the virtual text includes comments on target content in the first video based on a virtual character identity.
[0172] According to one or more embodiments of the present disclosure, playing the first video includes: obtaining video stream data; obtaining the first video based on the video stream data, and playing the first video in a first window on a video playback page; playing the second video includes: obtaining the second video based on the video stream data, and playing the second video in a second window on the video playback page; where the first window and the second window have a layer relationship, and the layer relationship is determined based on video playback states corresponding to the first window and the second window.
[0173] According to one or more embodiments of the present disclosure, playing the second video includes: obtaining, through the video stream data, time information corresponding to the first video, where the time information is used to characterize a time period corresponding to at least one content paragraph in the first video; and playing the second video according to the time information.
[0174] According to one or more embodiments of the present disclosure, playing the second video includes: parsing the video content currently played in the first video; and playing the second video according to the video content so that the voice speeches in the second video and the first video are not played simultaneously.
[0175] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an information generation device, including:
[0176] An acquisition module configured to acquire description data corresponding to a first video, where the description data represents historical comments of the first video;
[0177] A processing module configured to process the description data through a language model to generate virtual text, where the content of the virtual text includes comments on target content in the first video based on the identity of a virtual character;
[0178] A generation module configured to create a virtual object that reads the virtual text aloud and generate a second video based on the virtual object.
[0179] According to one or more embodiments of the present disclosure, the processing module is specifically configured to: obtain a corresponding target prompt word according to the identification information of the first video, where the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language; and use the language model to process the target prompt word and the description data to generate virtual text.
[0180] According to one or more embodiments of the present disclosure, the target prompt word includes character information, where the character information is used to describe the identity characteristics and / or preference characteristics of the virtual character; when the processing module uses the language model to process the target prompt word and the description data to generate virtual text, it is specifically configured to: use the language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
[0181] According to one or more embodiments of the present disclosure, the generation module is specifically configured to: create a virtual object; generate corresponding voice data according to the virtual text; drive the virtual object to perform corresponding actions and / or expressions based on the voice data, and save it as the second video.
[0182] According to one or more embodiments of the present disclosure, when creating a virtual object, the generation module is specifically configured to: generate feature information of the virtual object according to the video content of the first video, where the feature information is used to represent the appearance characteristics of the virtual object; and create the virtual object according to the feature information.
[0183] According to one or more embodiments of the present disclosure, when generating corresponding voice data according to the virtual text, the generating module is specifically configured to: obtain feature information of the virtual object, where the feature information is used to characterize the appearance features of the virtual object; and based on the feature information, process the virtual text to generate voice data having voice features matching the appearance features.
[0184] According to one or more embodiments of the present disclosure, the processing module is further configured to: obtain semantic features corresponding to the virtual text; and obtain corresponding video background materials according to the semantic features; when generating a second video based on the virtual object, the generating module is specifically configured to: generate a second video based on the virtual object and the video background materials.
[0185] According to one or more embodiments of the present disclosure, the description data includes at least one of the following: text, picture, video, and voice.
[0186] Fourthly, according to one or more embodiments of the present disclosure, an information display device is provided, including:
[0187] A first playback module, configured to play a first video;
[0188] A second playback module, configured to play a second video in response to a trigger instruction for the first video, where the second video includes a virtual object that reads the virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
[0189] According to one or more embodiments of the present disclosure, the first playback module is specifically configured to: obtain video stream data; obtain the first video based on the video stream data, and play the first video in a first window on a video playback page; when playing the second video, the second playback module is specifically configured to: obtain the second video based on the video stream data, and play the second video in a second window on the video playback page; where the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback states corresponding to the first window and the second window.
[0190] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module is specifically configured to: obtain, through the video stream data, time information corresponding to the first video, where the time information is used to characterize the time period corresponding to at least one content segment in the first video; and play the second video according to the time information.
[0191] According to one or more embodiments of the present disclosure, when playing the second video, the second playback module is specifically configured to: parse the video content currently played in the first video; and play the second video according to the video content, so that the voice in the second video and the voice in the first video are not played simultaneously.
[0192] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, including: at least one processor and a memory;
[0193] The memory stores computer-executable instructions;
[0194] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the information generation method described in the first aspect above and various possible designs of the first aspect, or executes the information display method described in the second aspect above and various possible designs of the second aspect.
[0195] In a fourth aspect, according to one or more embodiments of the present disclosure, there is provided a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the information generation method described in the first aspect above and various possible designs of the first aspect is implemented, or the information display method described in the second aspect above and various possible designs of the second aspect is implemented.
[0196] In a fifth aspect, according to one or more embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the information generation method described in the first aspect above and various possible designs of the first aspect is implemented, or the information display method described in the second aspect above and various possible designs of the second aspect is implemented.
[0197] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with technical features having similar functions disclosed in the present disclosure (but not limited to).
[0198] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0199] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. An information generation method, characterized in that, it includes: Obtain the description data corresponding to the first video, and the description data characterizes the historical comments of the first video; Process the description data through a language model to generate virtual text, and the content of the virtual text includes comments on the target content in the first video based on the identity of the virtual character; Create a virtual object that reads the virtual text aloud, and generate a second video based on the virtual object.
2. The method according to claim 1, characterized in that, The process of processing the description data through the language model to generate virtual text includes: Obtain the corresponding target prompt word according to the identification information of the first video, and the target prompt word is used to describe the evaluation requirements for at least one evaluation dimension based on natural language; Use the language model to process the target prompt word and the description data to generate virtual text.
3. The method according to claim 2, characterized in that, The target prompt word includes character information, and the character information is used to describe the identity characteristics and / or preference characteristics of the virtual character; The process of using the language model to process the target prompt word and the description data to generate virtual text includes: Use the language model to process the target prompt word and the description data to generate virtual text corresponding to the identity characteristics and / or preference characteristics.
4. The method according to claim 1, characterized in that, Creating a virtual object that reads the virtual text aloud and generating a second video based on the virtual object includes: Create a virtual object; Generate corresponding voice data according to the virtual text; Based on the voice data, drive the virtual object to perform corresponding actions and / or expressions, and save it as the second video.
5. The method according to claim 4, characterized in that, The creation of the virtual object includes: Generate the feature information of the virtual object according to the video content of the first video, and the feature information is used to characterize the appearance features of the virtual object; Create the virtual object according to the feature information.
6. The method according to claim 4, characterized in that, The generation of corresponding voice data according to the virtual text includes: Obtain the feature information of the virtual object, and the feature information is used to characterize the appearance features of the virtual object; Based on the feature information, process the virtual text to generate voice data with sound features matching the appearance features.
7. The method according to claim 1, characterized in that, The method further includes: Obtain the semantic features corresponding to the virtual text; Obtain the corresponding video background material according to the semantic features; The generation of the second video based on the virtual object includes: Generate a second video based on the virtual object and the video background material.
8. The method according to claim 1, characterized in that, The description data includes at least one of the following: Text, picture, video, voice.
9. An information display method, characterized in that, it includes: Play the first video; In response to a trigger instruction for the first video, a second video is played, wherein the second video includes a virtual object that reads virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
10. The method according to claim 9, wherein, the playing of the first video includes: obtaining video stream data; obtaining the first video based on the video stream data, and playing the first video in a first window within a video playback page; the playing of the second video includes: obtaining the second video based on the video stream data, and playing the second video in a second window within the video playback page; wherein, the first window and the second window have a layer relationship, and the layer relationship is determined based on the video playback states corresponding to the first window and the second window.
11. The method according to claim 9, wherein, the playing of the second video includes: obtaining, through the video stream data, time information corresponding to the first video, the time information being used to characterize the time period corresponding to at least one content segment in the first video; playing the second video according to the time information.
12. The method according to claim 9, wherein, the playing of the second video includes: analyzing the video content currently being played in the first video; playing the second video according to the video content, so that the voice speech in the second video and the voice speech in the first video are not played simultaneously.
13. An information generation device, wherein, comprises: an obtaining module, configured to obtain description data corresponding to a first video, the description data characterizing historical comments on the first video; a processing module, configured to process the description data through a language model to generate virtual text, the content of the virtual text including comments on the target content in the first video based on the virtual character identity; a generating module, configured to create a virtual object that reads the virtual text aloud, and generate a second video based on the virtual object.
14. An information display device, wherein, comprises: a first playing module, configured to play a first video; a second playing module, configured to play a second video in response to a trigger instruction for the first video, wherein the second video includes a virtual object that reads virtual text aloud, and the content of the virtual text includes comments on the target content in the first video based on the virtual character identity.
15. An electronic device, wherein, comprises: a processor and a memory; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory, so that the processor executes the information generation method according to any one of claims 1 to 8, or executes the information display method according to any one of claims 9 to 12.
16. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the information generation method described in any one of claims 1 to 8 is implemented, or the information display method described in any one of claims 9 to 12 is implemented.
17. A computer program product, comprising a computer program, wherein, when the computer program is executed by a processor, the information generation method described in any one of claims 1 to 8 is implemented, or the information display method described in any one of claims 9 to 12 is implemented.