Content presentation method and apparatus, device, and storage medium
Patent Information
- Application Number
- PCT/CN2026/078405
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-20
- Filing Date
- 2026-02-10
- Publication Date
- 2026-08-27
Smart Images

Figure CN2026078405_27082026_PF_FP_ABST
Abstract
Description
Content representation methods, devices, equipment and storage media
[0001] This application claims priority to Chinese Patent Application No. 202510192590.0, filed on February 20, 2025, entitled "Method, Apparatus, Device and Storage Medium for Content Presentation", the entire contents of which are incorporated herein by reference. Technical Field
[0002] The exemplary embodiments disclosed herein generally relate to the field of computers, and particularly to content presentation methods, apparatus, devices, computer-readable storage media, and computer program products. Background Technology
[0003] With the rapid development of internet technology, the internet has become an important platform for people to access and share content. Users can access the internet through terminal devices to obtain content. Moreover, with the development of information technology, the forms in which terminal devices or applications provide content to users are becoming increasingly diversified to meet users' diverse service needs. Summary of the Invention
[0004] In a first aspect of this disclosure, a content representation method is provided. The method includes: responding to a request for source text content; determining, based on analysis of the source text content, at least a first analysis result indicating summary information in the source text content and at least one second analysis result indicating at least one content unit of the source text content; obtaining role configurations for multiple speakers corresponding to the source text content; and generating first dialogue content for the multiple speakers based on the first analysis result, the at least one second analysis result, and the role configurations of the multiple speakers. The first dialogue content includes a first dialogue segment and at least one second dialogue segment following the first dialogue segment, the first dialogue segment being associated with summary information, and the at least one second dialogue segment being associated with at least one content unit.
[0005] In a second aspect of this disclosure, an apparatus for content representation is provided. The apparatus includes: a determining module configured to, in response to a request for source text content, determine, based on analysis of the source text content, at least a first analysis result indicating summary information in the source text content and at least a second analysis result indicating at least one content unit of the source text content; an acquiring module configured to acquire role configurations for a plurality of speakers corresponding to the source text content; and a generating module configured to generate first dialogue content for the plurality of speakers based on the first analysis result, the at least one second analysis result, and the role configurations of the plurality of speakers. The first dialogue content includes a first dialogue segment and at least one second dialogue segment following the first dialogue segment, the first dialogue segment being associated with summary information, and the at least one second dialogue segment being associated with at least one content unit.
[0006] In a third aspect of this disclosure, an electronic device is provided. The device includes at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor. When executed by the at least one processor, the instructions cause the device to perform the method of the first aspect.
[0007] In a fourth aspect of this disclosure, a computer-readable storage medium is provided. The computer-readable storage medium stores computer-executable instructions that can be executed by a processor to implement the method of the first aspect.
[0008] In a fifth aspect of this disclosure, a computer program product is provided, including computer-executable instructions, wherein when executed by a processor, the computer-executable instructions implement the method according to a first aspect of this disclosure.
[0009] It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0011] Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure may be implemented;
[0012] Figure 2 shows a flowchart illustrating the process of presenting some embodiments of the present disclosure;
[0013] Figures 3A and 3B respectively illustrate schematic diagrams of example processes described according to some embodiments of the present disclosure;
[0014] Figures 4A to 4C respectively show schematic diagrams of examples of user interfaces according to some embodiments of the present disclosure;
[0015] Figure 5 shows a schematic structural block diagram of an example device for content presentation according to some embodiments of the present disclosure; and
[0016] Figure 6 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure. Detailed Implementation
[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0018] In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". Other explicit and implicit definitions may also be included below.
[0019] In this document, unless explicitly stated otherwise, performing a step in response to A does not mean that the step is performed immediately after A, but may include one or more intermediate steps.
[0020] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0021] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.
[0022] For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information, thereby enabling the user to choose whether to provide personal information to the software or hardware such as electronic devices, applications, servers or storage media that perform the operation of the technical solution disclosed herein, based on the prompt message.
[0023] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, such as a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0025] As used in this paper, the term "model" refers to a model that learns the relationship between inputs and outputs from training data, enabling it to generate corresponding outputs for a given input after training. Model generation can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs using multiple layers of processing units. A neural network model is an example of a deep learning-based model. In this paper, "model" may also be referred to as a "machine learning model," "learning model," "machine learning network," or "learning network," and these terms are used interchangeably.
[0026] A neural network is a machine learning network based on deep learning. A neural network processes input and provides a corresponding output, typically consisting of an input layer, an output layer, and one or more hidden layers between the input and output layers. Neural networks used in deep learning applications often include many hidden layers, thus increasing the network's depth. The layers of a neural network are connected sequentially, so that the output of the previous layer is provided as the input to the next layer. The input layer receives the input to the neural network, while the output layer's output serves as the final output. Each layer of a neural network includes one or more nodes (also called processing nodes or neurons), each node processing the input from the layer above.
[0027] Machine learning typically comprises three phases: training, testing, and application (also known as inference). In the training phase, a given model is trained using a large amount of training data, iteratively updating its parameter values until the model can consistently generate inferences that meet the expected goals from the training data. Through training, the model can be considered to have learned the relationship between inputs and outputs (also known as the input-output mapping) from the training data. The parameter values of the trained model are determined. In the testing phase, test inputs are applied to the trained model to test whether it can provide the correct output, thus determining the model's performance. In the application phase, the model can be used to process actual inputs based on the trained parameter values to determine the corresponding output.
[0028] As mentioned above, with the rapid development of internet technology, the internet has become an important platform for people to access and share content. Users can access the internet through terminal devices to obtain content. Moreover, with the development of information technology, the forms in which terminal devices or applications provide content to users are becoming increasingly diversified to meet users' diverse service needs. For example, some application products offer podcast functionality, which can convert web page content into speeches by multiple speakers (also known as podcast scripts), and play the speeches in a dialogue format. However, traditional podcast functionality only converts web page content into podcast scripts in terms of format. Podcast scripts often lack logic and hierarchy, and the speaker roles are often one-dimensional and rigid, failing to resonate with listeners in many cases. Therefore, the playback quality of podcast functionality needs improvement.
[0029] In view of this, embodiments of the present disclosure propose an improved scheme for content representation. In this scheme, by analyzing source text content, a first analysis result and at least one second analysis result are determined. The first analysis result at least indicates summary information in the source text content, and the at least one second analysis result indicates at least one content unit of the source text content. In this scheme, the role configurations of multiple speakers corresponding to the source text content can also be obtained. Then, based on the first analysis result, the at least one second analysis result, and the role configurations of the multiple speakers, first dialogue content of multiple speakers is generated. The first dialogue content includes a first dialogue segment and at least one second dialogue segment following the first dialogue segment. The first dialogue segment is associated with summary information, and the at least one second dialogue segment is associated with at least one content unit.
[0030] In the embodiments of this disclosure, the main content and structure of the source text can be determined by analyzing the source text content. The dialogue content generated based on the analysis results includes a first dialogue segment and at least one second dialogue segment. The first dialogue segment can express the main content of the source text as a whole, while the at least one second dialogue segment can express the content related to each content unit of the source text, making the generated dialogue content clear and well-organized. Furthermore, by setting the role configuration of the speakers, the dialogue content can be made richer and more vivid, significantly improving the content quality of the dialogue content.
[0031] The following section provides a detailed description of various example implementations of this scheme, with reference to the accompanying drawings.
[0032] Example Environment
[0033] Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. In this example environment 100, an application 120 is installed on a terminal device 110. A user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110. For example, the application 120 can receive input from the user 140 via the terminal device 110, and the application 120 can also provide content corresponding to the input.
[0034] In some embodiments of this disclosure, application 120 can be any suitable application capable of presenting text content. For example, application 120 can present web pages, documents, or images containing text content. In some embodiments, application 120 can also be configured to output audio, video, or other modal content.
[0035] In some embodiments, if application 120 is active, terminal device 110 may display the user interface 150 of application 120. The user interface 150 may include various types of content that application 120 can provide, such as a dialogue page between the user and a digital assistant (in which the current dialogue and historical dialogues, including text dialogue content, may be displayed), a text content presentation interface, a voice playback interface, a video playback interface, and so on.
[0036] In some embodiments, application 120 or a digital assistant therein may utilize machine learning model 160 (which may include one or more machine learning models, such as machine learning model 160-1, machine learning model 160-2, ..., machine learning model 160-N, etc., where N is a positive integer. For ease of description, the one or more machine learning models are collectively referred to as machine learning model 160 herein) to support interaction with user 140. For example, application 120 or a digital assistant therein may utilize one or more machine learning models 160 to determine the content corresponding to the input of user 140.
[0037] Machine learning model 160 can be of different types. In some embodiments, one or more machine learning models 160 may be built based on a language model (LM). The machine learning model used is a content-generating model, capable of generating corresponding outputs based on model inputs. In some embodiments, the language model-based machine learning model can handle textual modal model inputs (e.g., natural language and / or machine language) and / or non-textual modal model inputs (e.g., images, speech, video, etc.), and can generate the desired output based on the model inputs and prompt words. Here, prompt words are used to guide the machine learning model to generate outputs that address the user needs indicated by the model inputs. In application scenarios supporting user dialogue, user 140's input can be provided to machine learning model 160 as at least a part of the model inputs (other parts may include prompt words). Based on the model outputs, content corresponding to user 140's inputs can be generated and provided to user 140.
[0038] In some embodiments, one or more machine learning models 160 may be speech-related models, including speech recognition (ASR) models and text-to-speech (TTS) models. The input to an ASR model is speech, and the output is text. The input to a TTS model is text, and the output is the corresponding speech.
[0039] In some embodiments, terminal device 110 communicates with server device 130 to provide services to application 120. As shown in FIG1, server device 130 may invoke machine learning model 160 to support human-computer dialogue between application 120 and user 140 based on the output of machine learning model 160. Terminal device 110 may be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, terminal device 110 may also support any type of user-facing interface (such as "wearable" circuitry). Server device 130 may be various types of computing systems / servers capable of providing computing power, including but not limited to mainframes, edge computing nodes, computing devices in cloud environments, etc. Server-side device 130 can be implemented, for example, based on a cloud environment.
[0040] It should be understood that the structure and function of the various elements in environment 100 are described for illustrative purposes only and do not imply any limitation on the scope of this disclosure.
[0041] Example process
[0042] The following description continues with reference to the accompanying drawings, outlining some exemplary embodiments of this disclosure. Figure 2 illustrates a flowchart of process 200 according to some embodiments of this disclosure. Part or all of process 200 may be implemented by terminal device 110 or by terminal device 110 in cooperation with other devices, such as by terminal device 110 in cooperation with server device 130. In the following description, for ease of discussion, the execution of process 200 will be described from the perspective of terminal device 110, but this is merely exemplary.
[0043] In box 210, in response to a request for source text content, terminal device 110 determines a first analysis result and at least one second analysis result based on analysis of the source text content. The source text content may include various text content related to terminal device 110 or application 120. In one example, the source text content may include text content presented by the user interface 150 of application 120, such as text content from a webpage, document, or image presented by user interface 150. In another example, the source text content may also include text content obtained through terminal device 110 or an attached device of terminal device 110.
[0044] In some embodiments of this disclosure, the first analysis result at least indicates summary information in the source text content. In some examples, the summary information may include at least one of the topic or summary content of the source text content, and the summary content may be a brief summary of the source text content. In some embodiments, the terminal device 110 may generate model input for the machine learning model 160-1 based on the source text content. The machine learning model 160-1 generates model input based on the model input. Then, based on the model output, the summary information indicating the source text content is determined. Thus, the machine learning model 160-1 can accurately determine the summary information of the source text content.
[0045] In some embodiments, terminal device 110 generates summary information of the source text content based on analysis of the source text content. Then, based on the summary information, terminal device 110 generates dialogue guidance content to obtain a first analysis result indicating the summary information and the dialogue guidance content. The dialogue guidance content is used to guide the generation of the initial dialogue segment (i.e., the first dialogue segment in the following text) in the dialogue content, thereby increasing the attractiveness of the initial dialogue segment. In some cases, the dialogue guidance content may also be referred to as an introduction.
[0046] As an example, Figure 3A illustrates a flowchart of an example process 300A for content presentation in some embodiments of this disclosure. As shown in Figure 3A, in block 304, terminal device 110 analyzes source text content 302 and determines summary information 306 of the source text content 302. Terminal device 110 can generate model input for machine learning model 160-1 based on the summary information 306. Using machine learning model 160-1, based on the model input, dialogue guidance content 308 is generated.
[0047] In some examples, terminal device 110 can generate dialogue guidance content based on summary information and reference guidance content indicating the style of the reference language. As an example, a database for storing reference quotations (i.e., reference guidance content) can be pre-built. Terminal device 110 can retrieve reference quotations from the database that match summary information 306 or source text content 302, and generate model input for machine learning model 160-1 based on summary information 306 and reference quotations. Then, using machine learning model 160-1, dialogue guidance content 308 is generated based on the model input.
[0048] In some embodiments of this disclosure, the at least one second analysis result indicates at least one content unit of the source text content. In some embodiments, the terminal device 110 may determine at least one source text fragment group based on the analysis of the source text content to form at least one content unit. Each source text fragment group includes at least one source text fragment from the source text content. Specifically, the terminal device 110 may split the source text content into several source text fragments based on the analysis of the source text content. Then, content-related source text fragments can be combined to form source text fragment groups. In this way, content-related source text fragments can be combined together, avoiding the dispersion and confusion of related content in the subsequently generated dialogue content, which is beneficial to improving the coherence and organization of the dialogue content. In some examples, each source text fragment group may also include a title indicating the main content described by the corresponding source text fragment group.
[0049] As an example, referring to Figure 3A, in box 310, terminal device 110 can generate model input for machine learning model 160-2 based on source text content 302. Machine learning model 160-2 analyzes the source text content 302 based on the model input to determine at least one group of source text fragments of the source text content 302, forming the at least one content unit 312.
[0050] In some examples, terminal device 110 can also determine the order of at least one content unit based on a predetermined logical relationship to obtain at least one second analysis result indicating the order of at least one content unit and the order between at least one content unit. The predetermined logical relationship can instruct the expression logic of subsequently generated text content regarding the at least one content unit. For example, the predetermined logical relationship can instruct the text content to express the content related to the at least one content unit according to expression logic such as chronological order, spatial order, event order, general-to-specific relationship, causal relationship, or sequential relationship. Determining the order of the at least one content unit helps improve the logicality of subsequently generated dialogue content.
[0051] As an example, terminal device 110 can determine the order of the at least one content unit based on the at least one content unit and the indication of a predetermined logical relationship using machine learning model 160-2. In this case, the second analysis result not only indicates the corresponding content unit, but also the order of the corresponding content units.
[0052] In some embodiments, the terminal device 110 may further determine a label indicating the expression type of at least one content unit, to be used as part of at least one second analysis result. The expression type indicates how a speaker expresses the corresponding content unit. In some examples, the expression type for a content unit may include narrative, argumentative (also known as discussion), explanatory, descriptive, lyrical, etc. Of course, the above expression types are merely exemplary, and any other appropriate expression type can be selected according to actual needs; the embodiments of this disclosure do not limit this. By determining the expression type of the content unit, the subsequently generated dialogue content can express the corresponding content unit in an appropriate manner, which is beneficial to improving the expressiveness and vividness of the dialogue content.
[0053] In some examples, terminal device 110 may generate at least one extended result based on content expansion of at least one content unit. Based on the at least one extended result, terminal device 110 determines a score indicating the dialectical characteristics of the at least one content unit. Then, based on the respective scores of the at least one content unit, terminal device 110 may determine a label indicating the expression type of the at least one content unit.
[0054] Each extended result may include at least one of the following: the topic described in the corresponding content unit, positive viewpoints on the topic, negative viewpoints on the topic, or a summary of the topic. A topic may indicate the subject or point of discussion described or discussed in the corresponding content unit. Positive viewpoints may include opinions that affirm or endorse the topic; they may also be referred to as positive or optimistic viewpoints. Negative viewpoints may include opinions that negate or question the topic; they may also be referred to as negative or pessimistic viewpoints. The summary may be a general description of the topic.
[0055] As an example, continuing with Figure 3A, the terminal device 110 can generate model inputs for the machine learning model 160-3 based on each of the at least one content unit 312. Using the machine learning model 160-3, the extended result 314 of each content unit 312 is determined based on each model input.
[0056] In some examples, dialectical characteristics can indicate the richness of the debate between positive and negative viewpoints in the expanded results. As an example, terminal device 110 can determine a score indicating the dialectical characteristics of at least one content unit based on the information density of the topic, positive viewpoints, negative viewpoints, and summary content in the expanded results. As another example, as shown in Figure 3A, terminal device 110 can generate model input for machine learning model 160-3 based on the expanded results 314. Using machine learning model 160-3, a score 324 is determined for each content unit based on the model input.
[0057] In some examples, a first threshold indicating the richness of debate can be predetermined. Terminal device 110 can compare the score to this first threshold. If the score exceeds the first threshold, it indicates that the richness of debate between the positive and negative viewpoints of the corresponding content unit is relatively high, making it suitable for expression through debate. Terminal device 110 can then determine the expression type of the corresponding content unit as a debate type and generate a label indicating the debate type. In this case, the corresponding second analysis result can include the content unit, the ranking of the content units, expanded results, and the label indicating the debate type.
[0058] If the score does not exceed the first threshold, it indicates that the richness of the debate between the positive and negative viewpoints of the corresponding content unit is relatively low, making it suitable to describe the corresponding content unit in a narrative manner. The terminal device 110 can determine the expression type of the corresponding content unit as a narrative type and generate a tag indicating the narrative type. In this case, the corresponding second analysis result may include the content unit, the ranking of the content units, and the second tag indicating the narrative type. It should be noted that the above method of determining the expression type is merely exemplary. In practical applications, any other suitable method can be used to determine the expression type; for example, the terminal device 110 may also use a machine learning model to determine the expression type of the content unit. The embodiments of this disclosure do not specifically limit this aspect.
[0059] In some embodiments, a second threshold (sometimes referred to herein as a rating threshold) below the first threshold may also be predetermined. The terminal device 110 can compare the rating of each content unit with the second threshold. If the rating of a content unit exceeds the second threshold, the terminal device 110 can retain the content unit. If the rating of a content unit does not exceed the second threshold, indicating that the richness and information density of the corresponding content unit are low and its impact on the integrity of the source text content is minimal, the terminal device 110 can remove the content unit corresponding to that rating. This ensures that subsequently generated dialogue content revolves around the main content, avoiding confusion in the dialogue.
[0060] Requests for source text content can be triggered in various ways. The following examples, illustrated with Figures 4A to 4C, illustrate how to trigger such requests, but should not be construed as being limited to these methods. In practical applications, any appropriate method can be chosen to trigger requests for source text content based on actual needs.
[0061] As an example, Figure 4A illustrates a schematic diagram of an example 400A of a user interface according to some embodiments of the present disclosure. In example 400A, the user interface 150 of application 120 may include a first area 402 for presenting web page content and an input field 404. If a trigger on the input field 404 is received (e.g., a single click, double click, or long press, etc.), the terminal device 110 determines that a request for web page content to be presented in the first area 402 has been received, and determines a first analysis result and at least one second analysis result based on the analysis of the web page content.
[0062] As another example, as shown in Figures 4A and 4B, in example 400A, the user interface 150 of application 120 presents a dialog entry 406 for a digital assistant. If a trigger is received on dialog entry 406, terminal device 110 presents a dialog interface 408 for the digital assistant in user interface 150. Dialogue interface 408 presents an input entry 410. If a trigger is received on input entry 410, terminal device 110 determines that a request for webpage content presented in first area 402 has been received, and based on analysis of the webpage content, determines a first analysis result and at least one second analysis result.
[0063] As another example, as shown in Figure 4C, with the dialog interface 408 already presented in the user interface 150, the user can input predetermined dialogue content 414 through the input field 412 in the dialog interface 408, such as "Help me generate a podcast from the webpage content and play it." The digital assistant can respond to this predetermined dialogue content by generating a request for the source text content.
[0064] In the return process 200, at box 220, the terminal device 110 obtains the role configurations of multiple speakers corresponding to the source text content. In some examples, the role configuration can indicate the master / servant role of the corresponding speaker in the subsequent dialogue content; for example, the role configuration can indicate that the speaker is a host or a guest. In other examples, the role configuration can support the personality attributes of the corresponding speaker. Of course, the above role configurations are only exemplary. In practical applications, the role of the speaker can be set from multiple dimensions, and the embodiments of this disclosure are not limited in this regard. In addition, it should be noted that the above speakers are actually virtual speakers, and the above role configurations do not involve personal privacy data. The purpose of obtaining the role configurations is to set the role attributes of the speakers so that the inventor's role attributes can match the source text content, thereby improving the vividness and attractiveness of the subsequently generated dialogue content.
[0065] In some embodiments, terminal device 110 can receive user instructions and determine the role configurations of multiple speakers based on those instructions. This allows users to configure speaker roles according to their personal preferences and actual needs, improving user satisfaction. As an example, FIG3B shows a flowchart of an example process 300B describing some embodiments of this disclosure. As shown in FIG3B, a set of speaker roles 330 can be pre-constructed. The role set 330 may include multiple predetermined speakers 332-1, 332-2, ..., 332-N. Users can use terminal device 110 or terminal device 110 to select a target speaker from the multiple predetermined speakers in the role set. Terminal device 110 can determine the role configurations of multiple target speakers based on the user's selection of the target speaker. Assuming the user selects predetermined speakers 332-1 and 332-2 as target speakers, terminal device 110 can obtain the role configuration 334-1 of predetermined speaker 332-1 and the role configuration 334-2 of predetermined speaker 332-2.
[0066] In some embodiments, terminal device 110 may also determine the role configurations of multiple inventors based on source text content or information related to the source text content. As an example, terminal device 110 may generate model inputs for machine learning model 160-4 based on source text content, summary information, or information related to each content unit. Machine learning model 160-4 is then used to determine the role configurations of multiple speakers.
[0067] Referring to Figure 2, in box 230, terminal device 110 generates first dialogue content for multiple speakers based on a first analysis result, at least one second analysis result, and the role configuration of multiple speakers. The first dialogue content includes a first dialogue segment and at least one second dialogue segment following the first dialogue segment. The first dialogue segment is associated with summary information, and the at least one second dialogue segment is associated with at least one content unit.
[0068] In some embodiments, a first dialogue segment can be used to initiate dialogue content. The first dialogue segment may include one or more dialogue statements related to summary information. As an example, the first dialogue segment may include one or more dialogue statements spoken solely by the main speaker (also known as the moderator). As another example, the first dialogue segment may include multiple dialogue statements spoken collaboratively by the main speaker and at least one co-speaker (also known as a guest).
[0069] In some embodiments, terminal device 110 can generate a first dialogue segment based on summary information, dialogue guidance content, and role configurations of multiple speakers. Combining summary information and dialogue guidance content can generate an expressive and engaging first dialogue segment, which is beneficial in attracting users to continue viewing or listening to subsequent dialogue content. As an example, as shown in FIG3B, terminal device 110 can obtain a first analysis result 336 and role configurations 334-1 and 334-2. The first analysis result 336 may include summary information and dialogue guidance content. Terminal device 110 can generate model input for machine learning model 160-5 based on the first analysis result 336 and role configurations 334-1 and 334-2. This model input is provided to machine learning model 160-5 to obtain the model output of machine learning model 160-5. Then, terminal device 110 can generate a first dialogue segment 338 based on this model output.
[0070] In some embodiments, the terminal device 110 can generate a second dialogue segment based on a second analysis result corresponding to the second dialogue segment, the role configuration of multiple speakers, and at least one preceding dialogue segment prior to the second dialogue segment. The second dialogue segment includes a transition statement and at least one dialogue statement used to transition from the at least one preceding dialogue segment to the second dialogue segment. The at least one dialogue statement is associated with a corresponding content unit. Specifically, the second dialogue segment generated in conjunction with at least one preceding dialogue segment includes a transition statement, which allows for a smooth transition from the preceding dialogue segment to the current second dialogue segment, thus improving the coherence of the generated dialogue content.
[0071] It is understood that the preceding dialogue segment may include a first dialogue segment or a second dialogue segment. As an example, as shown in Figure 3B, terminal device 110 can generate a second dialogue segment 342 based on the first dialogue segment 338, the second analysis result 340, and role configurations 334-1 and 334-2. Terminal device 110 can also generate a second dialogue segment 346 based on the second dialogue segment 342, the second analysis result 344, and role configurations 334-1 and 334-2. Then, the first dialogue content 348 is formed by the first dialogue segment 338, the second dialogue segment 342, and the second dialogue segment 346. It should be noted that although only the second dialogue segment 342 and the second dialogue segment 346 are shown in the example in Figure 3B, in actual applications, the dialogue content is not limited to including two second dialogue segments; the dialogue content may include one second dialogue segment, or it may include three or more second dialogue segments.
[0072] As can be seen from the content in section 220, the content contained in the second analysis result may also differ depending on the expression type. Therefore, the input content used to generate the second dialogue fragment for different expression types may also differ. The following explanation of the generation process of the second dialogue fragment, combining argumentation and narrative types, should be understood as merely illustrative.
[0073] In some examples, the expression type of a content unit may include a debate type. The second analysis result may include the corresponding content unit, the corresponding extended results, and a label indicating the debate type. Terminal device 110 can generate a second dialogue segment corresponding to the corresponding content unit based on the corresponding content unit, the corresponding extended results, and the label indicating the debate type. Since the extended results include topics, positive viewpoints, negative viewpoints, and summary content, combining the extended results to generate the corresponding second dialogue segment helps improve the debate richness of the second dialogue segment.
[0074] In other examples, the expression type of the content unit may include a narrative type. The second analysis result may include the corresponding content unit and a label indicating the narrative type. The terminal device 110 can generate a second dialogue fragment corresponding to the corresponding content unit based on the corresponding content unit and the label indicating the narrative type. In this way, the content described in the second dialogue fragment can be more closely similar to the original content unit.
[0075] In some embodiments, the first dialogue content may include at least one of first dialogue text or first dialogue audio. The first dialogue text may include dialogue text segments corresponding to multiple speakers, and the first dialogue audio may include multiple dialogue audio segments corresponding to multiple speakers. The terminal device 110 can present the first dialogue text and also play the first dialogue audio. In this way, it can not only present the dialogue text of the first dialogue content for user convenience, but also play the first dialogue content in a podcast-like format.
[0076] As an example, terminal device 110 can use machine learning model 160-5 to generate first dialogue text based on a first analysis result, at least one second analysis result, and the role configuration of multiple speakers. Then, terminal device 110 can also use a TTS model to generate first dialogue speech based on the first dialogue text.
[0077] As another example, referring to Figure 4C, terminal device 110 can present first dialogue text 416 through the array assistant's dialogue interface 408. Multiple speakers include a host and guests. The first dialogue text 416 includes dialogue text segments 420 and 424 corresponding to the host, and also includes dialogue text segment 422 corresponding to the guest. Simultaneously, terminal device 110 can also play first dialogue audio. The first dialogue audio may include audio segments of the host narrating dialogue text segments 420 and 424, and may also include audio segments of the guest narrating dialogue text segment 422.
[0078] In some embodiments, terminal device 110 can also acquire reference dialogue content indicating a reference language style. Then, based on the reference dialogue content, the first dialogue content is rewritten to obtain second dialogue content conforming to the reference language style. This enables the generation of second dialogue content conforming to a specific language style. In some examples, the reference language style may include, but is not limited to, colloquial language style, formal language style, dialectal language style, language style related to a specific person, etc. As an example, as shown in FIG3B, terminal device 110 can generate model input for machine learning model 160-5 based on reference dialogue content indicating a colloquial language style and first dialogue content 348. Using machine learning model 160-5, based on the model output, second dialogue content 350 conforming to a colloquial language style is generated to reduce the mechanical feel of the dialogue content.
[0079] It should be noted that the aforementioned machine learning models can be implemented as the same model or as different models. Furthermore, these machine learning models can be deployed on the terminal device 110 or on the server device 130. For example, the terminal device 110 can interact with the machine learning model 160 through the server device 130. The embodiments of this disclosure do not limit this aspect.
[0080] In this way, in the embodiments of this disclosure, it is possible to generate dialogue content that is clearly defined and hierarchical. Furthermore, by setting the roles of the speakers, the dialogue content can be made richer and more engaging, significantly improving the quality of the dialogue content.
[0081] Example devices and equipment
[0082] Embodiments of this disclosure also provide corresponding apparatus for implementing the methods or processes described above. Figure 5 shows a schematic structural block diagram of an example apparatus 500 for content representation according to certain embodiments of this disclosure. Apparatus 500 may be implemented as or included in terminal device 110. The various modules / components in apparatus 500 may be implemented by hardware, software, firmware, or any combination thereof.
[0083] As shown in Figure 5, the apparatus 500 includes: a determining module 510 configured to, in response to a request for source text content, determine, based on analysis of the source text content, at least a first analysis result indicating summary information in the source text content and at least a second analysis result indicating at least one content unit of the source text content; an acquiring module 520 configured to acquire the role configurations of multiple speakers corresponding to the source text content; and a generating module 530 configured to, based on the first analysis result, at least one second analysis result, and the role configurations of the multiple speakers, generate first dialogue content of the multiple speakers, the first dialogue content including a first dialogue segment and at least one second dialogue segment following the first dialogue segment, the first dialogue segment being associated with summary information, and the at least one second dialogue segment being associated with at least one content unit.
[0084] In some embodiments, the determining module 510 is further configured to: generate summary information of the source text content based on the analysis of the source text content; and generate dialogue guidance content based on the summary information to obtain a first analysis result indicating the summary information and the dialogue guidance content.
[0085] In some embodiments, the determining module 510 is further configured to generate dialogue guidance content based on summary information and reference guidance content indicating the style of the reference language.
[0086] In some embodiments, the determining module 510 is further configured to generate a first dialogue fragment based on summary information, dialogue guidance content, and the role configurations of multiple speakers.
[0087] In some embodiments, the determining module 510 is further configured to: determine at least one source text fragment group based on the analysis of the source text content to form at least one content unit, each source text fragment group including at least one source text fragment in the source text content; and determine the order of the at least one content unit based on a predetermined logical relationship to obtain at least one second analysis result indicating the order of the at least one content unit and the order between the at least one content unit.
[0088] In some embodiments, the determining module 510 is further configured to: determine a tag indicating the expression type of at least one content unit, as part of at least one second analysis result.
[0089] In some embodiments, the determining module 510 is further configured to: generate at least one extended result based on the content extension of at least one content unit, each extended result including at least one of the topic described by the corresponding content unit, a positive viewpoint on the topic, a negative viewpoint on the topic, or a summary of the topic; determine a score indicating the dialectical characteristics of at least one content unit based on the at least one extended result; and determine a label indicating the expression type of at least one content unit based on the respective scores of at least one content unit.
[0090] In some embodiments, the determining module 510 is further configured to: remove the content unit corresponding to the score in response to the fact that the score of the content unit in at least one content unit does not exceed the score threshold.
[0091] In some embodiments, the expression type includes a debate type, and the generation module 530 is further configured to: in response to a tag indicating that the expression type of the corresponding content unit is a debate type, generate a second dialogue fragment corresponding to the corresponding content unit based on the corresponding content unit, the extended result corresponding to the corresponding content unit, and the tag.
[0092] In some embodiments, the expression type includes a narrative type, and the generation module 530 is further configured to: in response to a tag indicating that the expression type of the corresponding content unit is a narrative type, generate a second dialogue fragment corresponding to the corresponding content unit based on the corresponding content unit and the tag.
[0093] In some embodiments, the generation module 530 is further configured to: generate a second dialogue segment based on a second analysis result corresponding to the second dialogue segment, the role configuration of multiple speakers, and at least one preceding dialogue segment prior to the second dialogue segment. The second dialogue segment includes a transition statement for transitioning from at least one preceding dialogue segment to the second dialogue segment and at least one dialogue statement associated with a corresponding content unit.
[0094] In some embodiments, the first dialogue content includes at least one of the first dialogue text or the first dialogue voice, and the device 500 further includes at least one of the following: a presentation module configured to present the first dialogue text, the first dialogue text including multiple dialogue text segments corresponding to multiple speakers, or a playback module configured to play the first dialogue voice, the first dialogue voice including multiple dialogue voice segments corresponding to multiple speakers.
[0095] In some embodiments, the apparatus 500 further includes: a rewriting module configured to acquire reference dialogue content indicating a reference language style; and to rewrite first dialogue content based on the reference dialogue content to obtain second dialogue content conforming to the reference language style.
[0096] The units and / or modules included in device 500 can be implemented in various ways, including software, hardware, firmware, or any combination thereof. In some embodiments, one or more units and / or modules can be implemented using software and / or firmware, such as machine-executable instructions stored on a storage medium. In addition to or as an alternative to machine-executable instructions, some or all of the units and / or modules in device 500 can be implemented at least partially by one or more hardware logic components. By way of example and not limitation, exemplary types of hardware logic components that can be used include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
[0097] Figure 6 shows a block diagram of an electronic device 610 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 610 shown in Figure 6 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 610 shown in Figure 6 may include or be implemented as the terminal device 110 of Figure 1 and the device 500 of Figure 5.
[0098] As shown in Figure 6, the electronic device 610 is in the form of a general-purpose electronic device. Components of the electronic device 610 may include, but are not limited to, one or more processors or processing units 610, memory 620, storage device 630, one or more communication units 640, one or more input devices 650, and one or more output devices 660. The processor 610 may be a physical or virtual processor and is capable of performing various processes according to executable instructions stored in memory 620. In a multiprocessor system, multiple processors execute computer-executable instructions in parallel to improve the parallel processing capability of the electronic device 610.
[0099] Electronic device 610 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 610, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 620 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 630 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data and can be accessed within electronic device 610.
[0100] Electronic device 610 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 6, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided. In these cases, each drive may be connected to a bus (not shown) via one or more data media interfaces. Memory 620 may include computer program product 625 having one or more executable instruction modules configured to perform various methods or actions of various embodiments of the present disclosure.
[0101] Communication unit 640 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 610 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 610 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
[0102] Input device 650 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 660 can be one or more output devices, such as a monitor, speaker, printer, etc. Electronic device 610 can also communicate with one or more external devices (not shown) via communication unit 640 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 610, or with any device that enables electronic device 610 to communicate with one or more other electronic devices (e.g., network card, modem, etc.). Such communication can be performed via input / output (I / O) interface (not shown).
[0103] According to an exemplary implementation of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to an exemplary implementation of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
[0104] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable and executable instructions.
[0105] These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processor of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-executable instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner; thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.
[0106] Computer-executable instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, executable instruction, or portion of instructions, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0108] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for content expression, comprising: responsive to a request for source text content, determining, based on an analysis of the source text content, a first analysis result indicative of at least summary information in the source text content and at least one second analysis result indicative of at least one content unit of the source text content; obtaining respective role configurations of a plurality of speakers corresponding to the source text content; and generating, based on the first analysis result, the at least one second analysis result and the role configurations of the plurality of speakers, first dialogue content of the plurality of speakers, the first dialogue content comprising a first dialogue segment and at least one second dialogue segment following the first dialogue segment, the first dialogue segment being associated with the summary information, and the at least one second dialogue segment being respectively associated with the at least one content unit. 2.The method of claim 1, wherein determining the first analysis result comprises: generating, based on the analysis of the source text content, the summary information of the source text content; and generating, based on the summary information, dialogue guide content to obtain the first analysis result indicative of the summary information and the dialogue guide content. 3.The method of claim 1 or 2, wherein generating the dialogue guide content comprises: generating, based on the summary information and reference guide content indicative of a reference language style, the dialogue guide content. 4.The method of any one of claims 1-3, wherein the first dialogue segment is determined by: generating, based on the summary information, the dialogue guide content and the role configurations of the plurality of speakers, the first dialogue segment. 5.The method of any one of claims 1-4, wherein determining the at least one second analysis result comprises: determining, based on the analysis of the source text content, at least one source text segment group to form the at least one content unit, each source text segment group comprising at least one source text segment in the source text content; and determining, based on a predetermined logical relationship, an order of the at least one content unit to obtain the at least one second analysis result indicative of the at least one content unit and the order among the at least one content unit. 6.The method of any one of claims 1-5, wherein determining the at least one second analysis result further comprises: determining a label indicative of an expression type of the at least one content unit to be respectively part of the at least one second analysis result. 7.The method of any one of claims 1-6, wherein determining the label indicative of the expression type of the at least one content unit comprises: generating, based on a content extension of the at least one content unit, at least one extension result, each extension result comprising at least one of a topic described by a corresponding content unit, a positive view on the topic, a negative view on the topic, or summary content on the topic; determining, based on the at least one extension result, a score indicative of a dialectical characteristic of the at least one content unit; and determine, based on the scores of the at least one content unit respectively, a label indicating an expression type of the at least one content unit. 8.The method of any one of claims 1-7, further comprising: in response to a score of a content unit in the at least one content unit not exceeding a score threshold, removing the content unit corresponding to the score. 9.The method of any one of claims 1-8, wherein the expression type comprises a debate type, and wherein the second dialogue segment is generated by: in response to the label indicating the expression type of a corresponding content unit as the debate type, generating the second dialogue segment corresponding to the corresponding content unit based on the corresponding content unit, an expansion result corresponding to the corresponding content unit, and the label. 10.The method of any one of claims 1-8, wherein the expression type comprises a narrative type, and wherein the second dialogue segment is generated by: in response to the label indicating the expression type of a corresponding content unit as the narrative type, generating the second dialogue segment corresponding to the corresponding content unit based on the corresponding content unit and the label. 11.The method of any one of claims 1-10, wherein the second dialogue segment is generated by: generating the second dialogue segment based on a second analysis result corresponding to the second dialogue segment, a role configuration of the plurality of speakers, and at least one preceding dialogue segment preceding the second dialogue segment, the second dialogue segment comprising a transition sentence for transitioning from the at least one preceding dialogue segment to the second dialogue segment and at least one dialogue sentence associated with a corresponding content unit. 12.The method of any one of claims 1-11, wherein the first dialogue content comprises at least one of a first dialogue text or a first dialogue speech, and the method further comprises at least one of: presenting the first dialogue text, the first dialogue text comprising a plurality of dialogue text segments corresponding to the plurality of speakers, or playing the first dialogue speech, the first dialogue speech comprising a plurality of dialogue speech segments corresponding to the plurality of speakers. 13.The method of any one of claims 1-12, further comprising: obtaining reference dialogue content indicating a reference language style; and rewriting the first dialogue content based on the reference dialogue content to obtain the second dialogue content conforming to the reference language style. 14.An apparatus for content representation, comprising: a determination module configured to, in response to a request for source text content, determine, based on an analysis of the source text content, a first analysis result indicating at least summary information in the source text content and at least one second analysis result indicating at least one content unit of the source text content; an obtaining module configured to obtain a role configuration of each of a plurality of speakers corresponding to the source text content; and a generation module configured to generate the second dialogue content based on the first analysis result, the at least one second analysis result, and the role configuration of each of the plurality of speakers. The generating module is configured to generate, based on the first analysis result, the at least one second analysis result, and roles of the plurality of speakers, first conversation content of the plurality of speakers, the first conversation content including a first conversation segment and at least one second conversation segment located after the first conversation segment, the first conversation segment being associated with the summary information, and the at least one second conversation segment being respectively associated with the at least one content unit.
15. An electronic device, comprising: at least one processor; and at least one memory coupled to the at least one processor and storing instructions for execution by the at least one processor, the instructions, when executed by the at least one processor, cause the electronic device to perform the method according to any one of claims 1-13.
16. A computer-readable storage medium having computer-executable instructions stored thereon that are executable by a processor to implement the method according to any one of claims 1-13.
17. A computer program product comprising computer-executable instructions that, when executed by a processor, implement the method according to any one of claims 1-13.