Target text processing method, device, equipment and medium

By using the enhanced constraint information and segmentation processing methods of the target domain in the large language model to generate multi-level directories and summary content, the problems of insufficient accuracy and richness of target text extraction in conference listening scenarios are solved, and more efficient text summary generation is achieved.

CN119719360BActive Publication Date: 2025-09-23BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411775429.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-09-23
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

In conference recording scenarios, existing technologies have difficulty effectively extracting the subject content from the target text, especially in long text and multi-domain scenarios. There are auditory hallucinations and difficulty in understanding, resulting in insufficient extraction accuracy and richness.

Method used

By obtaining the enhanced constraint information of the target domain to generate target prompt information, combined with the large language model, the focus on the target text is enhanced, and the segmented processing and multi-level directory generation methods are used to generate the target domain outline and digital clustering summary content, reduce irrelevant content, and improve the extraction accuracy and richness.

Benefits of technology

It achieves accurate extraction of target text in multi-domain scenarios, reduces auditory hallucinations and difficulty in understanding, improves the accuracy and pertinence of text summary content, and the generated summary content is more in line with the needs of the target field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119719360B_ABST
    Figure CN119719360B_ABST
Patent Text Reader

Abstract

The present disclosure provides a target text processing method, apparatus, device and medium, which relate to the field of data processing, specifically to the fields of speech recognition, human-computer interaction, artificial intelligence and large model technology. The specific implementation scheme is as follows: obtaining the target text; obtaining the target prompt information corresponding to the target field; the target prompt information is generated according to the enhanced constraint information corresponding to the enhanced type in the target field; the target prompt information corresponding to the target field and the target text are input into the large language model for processing, and the text summary content is output. The target prompt information is used by the large language model to increase the attention of the enhanced constraint information to enhance the corresponding content in the text summary content. The embodiment of the present disclosure can enhance the richness and accuracy of the information extracted from the target text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, specifically to the fields of speech recognition, human-computer interaction, artificial intelligence and large model technology, and in particular to a target text processing method, device, equipment and medium. Background Art

[0002] In the meeting recording scenario, the meeting topic content can be extracted from the speech-converted text.

[0003] Large Language Model (LLM) can be used to extract topic content from target text. Summary of the Invention

[0004] The present disclosure provides a target text processing method, apparatus, device and medium.

[0005] According to one aspect of the present disclosure, a target text processing method is provided, comprising:

[0006] Get the target text;

[0007] Obtaining target prompt information corresponding to the target field; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target field;

[0008] The target prompt information corresponding to the target field and the target text are input into the large language model for processing, and a text summary content is output. The target prompt information is used by the large language model to increase attention to the strengthened constraint information to enhance the corresponding content in the text summary content.

[0009] According to one aspect of the present disclosure, there is provided a target text processing apparatus, comprising:

[0010] A target text acquisition module is used to acquire the target text;

[0011] A prompt information determination module is used to obtain target prompt information corresponding to the target field; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target field;

[0012] The text summary extraction module is used to input the target prompt information corresponding to the target field and the target text into the large language model for processing and output text summary content. The target prompt information is used by the large language model to increase attention to the strengthened constraint information to enhance the corresponding content in the text summary content.

[0013] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the target text processing method described in any embodiment of the present disclosure.

[0017] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the target text processing method described in any embodiment of the present disclosure.

[0018] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the target text processing method described in any embodiment of the present disclosure when executed by a processor.

[0019] The embodiments of the present disclosure can enhance the richness and accuracy of information extraction from the target text.

[0020] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0022] Figure 1 is a flowchart of a target text processing method disclosed in an embodiment of the present disclosure;

[0023] Figure 2 is a flowchart of another target text processing method disclosed in an embodiment of the present disclosure;

[0024] Figure 3 is a flowchart of another target text processing method disclosed in an embodiment of the present disclosure;

[0025] Figure 4 is a schematic diagram of a page of text summary content disclosed according to an embodiment of the present disclosure;

[0026] Figure 5 is a structural diagram of a target text processing device disclosed in an embodiment of the present disclosure;

[0027] Figure 6 It is a block diagram of an electronic device according to the target text processing method disclosed in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0029] Figure 1 This is a flowchart of a target text processing method according to an embodiment of the present disclosure. This embodiment is applicable to extracting key content from a target text. This method can be performed by a target text processing device, which can be implemented using software and / or hardware and is specifically configured in an electronic device with certain data processing capabilities, such as a server device.

[0030] S101. Obtain target text.

[0031] The target text can be a meeting transcript that records the content of a meeting. Voice is captured and recognized to obtain the target text. The voice can refer to the speech from a meeting. Meeting voice can be captured and the target text obtained through voice recognition. The target text typically includes attribute information and key information about the meeting. Attribute information about the meeting can include the meeting time, location, participants, and meeting title (meeting theme). Key information about the meeting can refer to information extracted from the substantive content of the meeting.

[0032] S102: Acquire target prompt information corresponding to the target domain; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target domain.

[0033] Among them, the target field may refer to the field to which the voice content to be processed belongs. In some embodiments, the target field is news, and the target text should be a conference text in the news field. Exemplarily, the target fields may include: finance, education, entertainment, transportation, construction, environmental protection, energy, information technology, medical care and media, etc. The target prompt information is used to prompt how to process the target text. The target prompt information may refer to the description information of the processing task for processing the target text. In some embodiments, a large language model may be used to process the target text, and the target text and the target prompt information are fused and input into the large language model for processing. In some embodiments, the target prompt information is used to prompt the large language model how to process the target text. The target prompt information may include a prompt template prompt.

[0034] Among them, the enhancement type may refer to the dimension of data enhancement. The enhancement constraint information corresponds to the enhancement type. The enhancement constraint information may refer to the descriptive information corresponding to the enhancement type. In some embodiments, the enhancement constraint information may impose enhanced constraints on the output content, enhance the text comprehension ability, and enhance the ability to extract the content of the target domain. At least one enhancement type may be preset for the target domain, and the enhancement types corresponding to different domains are independent of each other. The enhancement constraint information corresponding to different enhancement types is independent of each other. In some embodiments, the target prompt information generated by the enhancement constraint information is used to provide to the large language model, and accordingly, the enhancement constraint information may be natural language content.

[0035] In some embodiments, generating target prompt information based on enhanced constraint information may be performed by obtaining standard prompt information, adding the enhanced constraint information to the standard prompt information, and obtaining the target prompt information.

[0036] S103: Input the target prompt information corresponding to the target field and the target text into the large language model for processing, and output text summary content. The target prompt information is used by the large language model to increase attention to the strengthened constraint information to enhance the corresponding content in the text summary content.

[0037] The text summary content may refer to the key and concise content representing the target text. In some embodiments, the text summary content may include target domain outline content, digital cluster summary content, and text summary content. The target prompt information and the target text are fused to obtain input data, which is then input into the large language model for processing and outputting the text summary content. In some embodiments, the target prompt information and the target text are input into the large language model, which encodes the input data to obtain an encoding result, and the large language model decodes the encoding result to obtain the text summary content.

[0038] In practice, strengthening constraint information can cause the large language model to pay more attention to the content of the strengthened constraint information, thereby allowing the large language model to impose strong constraints on the content of the text summary, extracting precise content from the target text to enhance the speech content, etc. Strengthening the constraint information to be content related to the target domain can make the large language model pay more attention to content related to the target domain and reduce content irrelevant to the target domain, thereby enhancing the target domain content in the text summary generated by the large language model.

[0039] For example, the enhanced constraint information is used to constrain the output format of the text summary content, so that the text summary content output by the large language model is the content of the output format specified in the enhanced constraint information.

[0040] For another example, the enhanced constraint information is used to limit the content extracted from the target text, so that the text summary content output by the large language model includes the content specified in the enhanced constraint information.

[0041] For another example, strengthening constraint information is used to enhance the accuracy of the content extracted from the target text, so that the large language model enhances the accuracy of the content extracted from the text summary content.

[0042] According to the technical solution disclosed herein, target prompt information is generated based on the reinforcement constraint information corresponding to the reinforcement type of the target domain, so that the large language model pays more attention to the reinforcement constraint information, so that the large language model can perform data enhancement on the text summary content, reduce auditory hallucinations and comprehension difficulty in the text summary content, reduce content irrelevant to the target domain in the text summary content, increase content related to the target domain in the text summary content, and improve the accuracy of the text summary content in the target domain.

[0043] In an optional embodiment, the target prompt information corresponding to the target field and the target text are input into a large language model for processing, and text summary content is output, including at least one of the following: the outline prompt information corresponding to the target field and the target text are input into a large language model for processing to generate the target field outline content; and the digital prompt information corresponding to the target field and the target text are input into a large language model for processing to generate the digital cluster summary content.

[0044] Among them, the target domain outline content may refer to the outline content adapted to the target domain. The digital cluster summary content may refer to the summary content of pure data in the target domain. The outline prompt information is used to prompt how to extract the target domain outline content from the target text. The digital prompt information is used to prompt how to extract the digital cluster summary content from the target text. In fact, some fields, such as finance, information technology, transportation or education, involve data statistical analysis. In these fields involving data analysis, data is the representative key content. The target domain outline content can be understood as the summary content extracted from the target text from the perspective of content. Exemplarily, the target domain outline content may include at least one directory, and the original text in the target text corresponding to each directory. The digital cluster summary content may extract the summary content from the target text from the perspective of data. Exemplarily, the digital cluster summary content may include at least one data class, and the attribute information and attribute values ​​of the data included in each data class.

[0045] In one example, the target field is finance, and the text summary content includes target field outline content and digital clustering summary content.

[0046] It can be seen that by configuring the text summary content as the target field outline content and / or digital clustering summary content, it is possible to obtain text summary content from multiple dimensions, flexibly increase the text summary content for the target field, enrich the text summary content, and improve the pertinence of the text content.

[0047] Figure 2 It is a flowchart of another target text processing method disclosed in accordance with an embodiment of the present disclosure, which is further optimized and expanded based on the above technical solution and can be combined with the above various optional implementation methods. The target prompt information corresponding to the target field and the target text are input into the large language model for processing, and the text summary content is output, which is specifically as follows: the outline prompt information corresponding to the target field and the target text are input into the large language model for processing to generate the target field outline content; and the digital prompt information corresponding to the target field and the target text are input into the large language model for processing to generate the digital cluster summary content; and when the enhanced type includes: when segmented processing, the outline prompt information corresponding to the target field and the target text are input into the large language model for processing to generate the target field outline content, which is specifically as follows: according to the first prompt information and the target text, a multi-level directory is generated; the multi-level directory is divided into tasks to generate at least one directory information extraction task; for each directory information extraction task, the target text corresponding to the directory of the directory information extraction task is generated according to the directory information extraction task, the second prompt information and the target text.

[0048] S201: Obtain target text.

[0049] S202: Obtain target prompt information corresponding to the target domain; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target domain, and the reinforcement type includes: segmented processing.

[0050] Enhanced segmentation can involve segmenting the target text during outline generation. In practice, large language models limit the output text length. When the target text is too long, requiring a longer output text that even exceeds the limit, large language models often experience auditory hallucinations, incorrect output formatting, and long sentences without punctuation.

[0051] Through the enhanced type of segmentation processing, the target text can be segmented and integrated, thereby avoiding the reduction in output accuracy caused by the output length limitation of the large language model.

[0052] S203: Generate a multi-level directory according to the first prompt information and the target text.

[0053] The first prompt information is used to generate a directory. A multi-level directory may refer to a directory with at least two levels. A multi-level directory includes multiple directories. Directories at adjacent levels have a subordinate relationship. Optionally, the multi-level directory includes two levels of directories, wherein the first level directory may include at least one second level directory, and the first level directory may not include the second level directory. Optionally, the multi-level directory includes a third level directory, wherein the second level directory may include at least one third level directory, and the second level directory may not include the third level directory.

[0054] In some embodiments, the first prompt information is combined with the target text to obtain input data, which is then fed into a large language model to output a multi-level directory. For example, the first prompt information is: Please output an outline directory of the target text.

[0055] In an example, the outline directory structure is as follows:

[0056] #1 Level Directory

[0057] ##Level 2 subdirectory

[0058] ##Level 2 subdirectory

[0059] #1 Level Directory

[0060] ##Level 2 subdirectory

[0061] ##Level 2 subdirectory

[0062] ##Level 2 subdirectory

[0063] S204: Divide the multi-level directory into tasks to generate at least one directory information extraction task.

[0064] Among them, task division can refer to processing multiple directories separately. The directory information extraction task can refer to the task of mapping the original text of the target text for the corresponding directory. In fact, different directories usually correspond to different paragraphs in the target text. Extracting the original text of a part of the directory is actually extracting the original text of some paragraphs in the target text, which is equivalent to extracting the target text in segments. It should be noted that different directories do not include directories with a subordinate relationship. For example, a first-level directory includes a second-level directory. The first-level directory is usually the integration result of the original text corresponding to the included second-level directory. Accordingly, the original text corresponding to the first-level directory and the original text corresponding to the included second-level directory have repeated content. The multi-level directory can be divided according to the hierarchical structure in the multi-level directory and the number of directories at the same level. The directories divided into the same task generate a directory information extraction task and serve as the directory corresponding to the directory information extraction task.

[0065] In some embodiments, directories with a subordinate relationship are grouped into the same task. If a primary directory includes a secondary directory, the secondary directory is actually a directory derived by dividing and summarizing the original text corresponding to the primary directory. Therefore, the original text corresponding to the primary directory and the original text corresponding to the included secondary directory may contain duplicate content. To avoid repeated and redundant operations, directories with a subordinate relationship are grouped into the same directory information extraction task for processing.

[0066] S205 . For each of the directory information extraction tasks, generate associated content corresponding to the directory of the directory information extraction task according to the directory information extraction task, the second prompt information, and the target text.

[0067] Among them, the second prompt information is used to extract the text corresponding to the directory from the target text. The associated content of the directory can refer to the extracted content of the text corresponding to the directory in the target text, wherein the corresponding text in the target text includes at least one sentence. In fact, the associated content can be the content obtained by further subdividing the directory. The associated content is different from the original text of the target text. The associated content is the summarized and refined content. The corresponding directory, the second prompt information and the target text in the directory information extraction task are integrated to obtain input data, and the input data is input into the large language model to output the associated content corresponding to each directory in the directory information extraction task. Exemplarily, the second prompt information is: Please output the original text corresponding to the outline directory of the target text.

[0068] In the hierarchical directory of affiliation, the original text corresponding to the higher-level directory is further summarized in sections to obtain the included lower-level directories. Accordingly, the target text content can be extracted only from the lowest-level directory, and the higher-level directories can be integrated layer by layer to obtain their corresponding related content. Different directories without affiliation have different related content, and the directory information extraction tasks are independent of each other, so different directory information extraction tasks can be executed in parallel.

[0069] S206, generating target domain outline content according to the associated content corresponding to the directory of each directory information extraction task; and / or

[0070] The target domain outline content can be generated from all generated directories and the associated content corresponding to each directory. The target domain outline content retains the directory structure, and there is a correspondence between the associated content and the directory. This allows the directory to be extracted and filled from the continuous content of the target text, making the directory more focused on a small range of associated content in the target text.

[0071] The embodiment of the present disclosure chooses to generate an outline rather than directly generate a report because the large language model has insufficient output capabilities for long texts. A single long text input cannot output content that meets the requirements. However, the content extracted from different directory structures in the outline can be extracted separately, which can fully utilize the understanding and generation capabilities of the large language model and split the outline into multiple directory structures to output accurate parts.

[0072] Furthermore, for scenarios where the target domain is finance, we choose to generate an outline based on the document structure rather than directly generating financial topics. This takes into account the non-orthogonal nature of the topical meanings. For example, while the focus of content varies between company performance and financial status, industry analysis, and development trends, the scope of coverage is generally similar. If the large language model generates both company performance and financial status topics simultaneously, the subsequent processing logic will often extract identical content to fill in the gaps. This will not only make the financial report lack logical structure, but also extract a large amount of duplicate content, reducing the accuracy of the extracted related content.

[0073] S207: Input the digital prompt information corresponding to the target field and the target text into a large language model for processing to generate digital cluster summary content.

[0074] According to the technical solution disclosed in the present invention, by generating a multi-level directory from the target text, the hierarchy of the directory can be increased, the situation where general topics cause the inability to distinguish different texts, and the situation where text duplication is caused by directory-based extraction can be avoided. By increasing the task of dividing the directory for information extraction, the extracted information can be made more focused, and the outline can be associated with information extracted from continuous directory texts to make the outline structure more reasonable and accurate. By dividing multiple tasks, it can be avoided that the output text exceeds the word count and the accurate content cannot be extracted.

[0075] In an optional embodiment, the multi-level directory is divided into tasks to generate at least one directory information extraction task, including: in the multi-level directory, obtaining multiple high-level directories; dividing each of the high-level directories into at least one task, with the upper limit of the number of divided tasks being a specified number; for each high-level directory, dividing the low-level directories associated with the high-level directory into the tasks divided by the high-level directory, to generate at least one directory information extraction task.

[0076] High-level directories include low-level directories. Too many tasks waste resources, while too few and a long target text can reduce the accuracy of the associated content, especially given the output text length limit. Therefore, a reasonable number of tasks should be set. The specified number can be determined based on experimentation. For example, four tasks are used.

[0077] In some embodiments, when the number of the highest-level directories is less than a specified number, the number of the highest-level directories can be used as the number of divided tasks; when the number of the highest-level directories is greater than or equal to the specified number, the specified number can be used as the number of divided tasks. In some embodiments, when the length of the target text is short, a smaller number of tasks can be generated, for example, 1 task, 2 tasks, or 3 tasks. The number of divided tasks can be determined based on the length of the target text. For example, when the length of the target text is greater than or equal to a preset length threshold, the number of divided tasks is determined to be a maximum of 4. For another example, based on the length of the target text, the length range can be determined from a plurality of preset length ranges corresponding to the number of tasks, and the corresponding number of tasks can be determined as the number of divided tasks.

[0078] Among them, the low-level directory associated with a high-level directory can refer to the directory subordinate to the high-level directory. Typically, the target text generates at least one high-level directory. For each high-level directory, the content corresponding to the high-level directory is subdivided to generate at least one low-level directory. The generated low-level directory is subordinate to the high-level directory and is associated with the high-level directory. Directly dividing the high-level directory and assigning the low-level directories included in the high-level directory to the same task of the high-level directory can extract continuous content, avoiding the duplication of related content caused by different tasks targeting the same continuous content, which can easily lead to extraction errors.

[0079] It can be seen that by dividing the tasks of high-level directories, generating directory information extraction tasks, and dividing the low-level directories associated with the high-level directories into the tasks where the high-level directories are located, the task division operation can be simplified, and at the same time, the extraction of the same continuous content can be focused on the same task, thereby improving the accuracy of the associated content and generating higher-quality outline content.

[0080] In an optional embodiment, dividing each of the high-level directories into at least one task includes: obtaining the directory order between each of the high-level directories; dividing each of the high-level directories into at least one task according to the directory order, wherein the number of directories included in the task of the high-level directory with an earlier directory order is less than the number of directories included in the task of the high-level directory with a later directory order.

[0081] The directory order can represent the order of the corresponding content in the target document. When displaying high-level directories, they are typically arranged in the order of their content in the target document. High-level directories can be evenly divided among a specified number of tasks, with different tasks containing different high-level directories. If even division is not possible, some tasks may contain more directories than others.

[0082] Generally speaking, the content in the first half of the target text is more important than the content in the second half. For content extraction tasks, large language models with long text output capabilities and product formats tend to experience a decrease in extraction accuracy as the amount of content they process increases. By ensuring that tasks with higher-level directories that come first include fewer directories than tasks with higher-level directories that come later, the large language model can focus more on the first half of the target text and extract higher-quality content.

[0083] The principle of task division is based on the number of first-level directories generated. A subsequent second round of 1-4 LLM tasks is performed based on the number of first-level directories generated. If more than four first-level directories are generated, which is the most common case, the first-level directories (combined with the included second-level directories) are evenly distributed among the four subsequent tasks. Considering the long text output capability of the large language model and the product form, the number of generated directories is small at the beginning and large at the end.

[0084] It can be seen that by making the number of directories included in the tasks of the high-level directory with an earlier directory order smaller than the number of directories included in the tasks of the high-level directory with a later directory order, the former tasks can be divided into fewer directories and the latter tasks can be divided into more tasks, so that the content extracted from the directory with an earlier directory order has higher quality and sufficient content extraction can be achieved.

[0085] In an optional embodiment, the multi-level directory includes: a first-level directory and a second-level directory.

[0086] Among them, the first-level directory can be a high-level directory, and the second-level directory can be a low-level directory. Or the first-level directory can be a low-level directory, and the second-level directory can be a high-level directory.

[0087] In an example, the extracted multi-level directory is as follows:

[0088] #Company Financial Performance

[0089] First Quarter Financial Highlights

[0090] ##Revenue growth and gross profit margin improvement

[0091] ##Cost Control and Operational Efficiency

[0092] The first round involves generating a multi-level table of contents. The second round involves extracting the associated content within the table of contents (i.e., extracting information from the table of contents). If the first round only extracts the first-level table of contents, such as "Company Financial Performance," the extracted table of contents is too general, making it difficult for the current LLM task to extract specific content from a specific text. This results in significant confusion in information extraction across different topics, and a significant amount of information may be duplicated for similar topics. Conversely, if the extracted table of contents is too detailed, for example, retaining a 3-4-level table of contents structure, the LLM's comprehension becomes more difficult, especially if the LLM's understanding differs between the two rounds. If the second-round LLM fails to recognize a sub-level extracted in the first round, it can easily experience auditory hallucinations during the information extraction process. This means that the LLM uses its existing knowledge to express an opinion on a specific topic, rather than being faithful to the original text. Whether or not the extraction is faithful to the original text is difficult to avoid through the current LLM's instruction adjustment. Experiments have shown that retaining the first and second levels of the table of contents is the optimal option.

[0093] It can be seen that by adopting the first and second level directories as the outline structure, we can avoid general topics that cannot distinguish different texts, resulting in repeated extracted related content, and avoid auditory hallucinations and associations caused by detailed topics, resulting in inaccurate related content. Therefore, adopting the first and second level directories can improve the accuracy of the outline content.

[0094] In an optional embodiment, the enhancement type includes: character enhancement processing; the outline prompt information corresponding to the target field is generated by: obtaining the general prompt information corresponding to the target field; obtaining the enhancement prompt information corresponding to the target field; generating the outline prompt information corresponding to the target field based on the general prompt information and the enhancement prompt information.

[0095] Among them, the reinforcement type of the role reinforcement processing may refer to the type of strengthening the large language model's understanding of its own role when generating the target domain outline content. Outline prompt information may refer to the outline content related to the target domain. General prompt information may refer to prompt information that is available in any field. Strengthened prompt information may refer to prompt information dedicated to the target domain. The strengthened prompt information for different target domains is different. In some embodiments, the general prompt information may be prompt information of prompt (prompt template). Strengthened prompt information may be prompt information of system. Among them, system is used to set the behavior, role and background of the large language model. Usually, system is often used to start a conversation, give a general direction of a conversation, or set the tone and style of the conversation. It can help set the context of the conversation so that the large language model can better understand its role in the conversation. Add the strengthened prompt information system to the general prompt information prompt to generate the target prompt information.

[0096] In some embodiments, outline prompt information can be generated in real time during human-computer interaction, or outline prompt information can be generated in advance. During human-computer interaction, when character enhancement is required, the outline prompt information with enhanced prompt information added is directly used.

[0097] It can be seen that by adding enhanced prompt information to the general prompt information, the correlation between the output of the large language model and the target domain can be strongly constrained, the domain relevance of the target domain in the outline content can be improved, and redundant content irrelevant to the target domain can be reduced.

[0098] In an optional embodiment, the acquiring of the general prompt information corresponding to the target domain includes: acquiring an extended domain corresponding to the target domain; and generating the general prompt information corresponding to the target domain according to the description information of the target domain and the description information of the extended domain.

[0099] The extended domain may refer to a domain that is related to the target domain but outside the target domain. The extended domain may be considered as a supplementary extension of the target domain. The descriptive information may refer to the content describing the domain. For example, the descriptive information may be an example of the domain.

[0100] In some embodiments, the target domain is finance. For example, common financial topics may include company profile, major products, industry analysis, development trends, competitive advantages, and team status. Extended domain topics may include company, business, industry, finance, and product information. The target domain and extended domains are added to the general prompt information to constrain the topic of the target text to be relevant to the target domain.

[0101] In one example, the large language model determines that the target text does not belong to the target domain and the extended domain based on general prompt information, does not generate the target domain outline content for the target text, and can output content that the target text does not belong to the target domain.

[0102] For example, based on general prompt information, the large language model determines that some content in the target text does not belong to the target domain and the extended domain. This part of the content does not generate the target domain outline content, and the final output target domain outline content does not include this part of the content and related information.

[0103] In reality, conference texts, especially those for voice conferences, include opening remarks, participant introductions, and small talk questions and answers. The large language model generates an outline structure for this content, but this is actually irrelevant to the target domain and is not the focus of the target domain. The extraction of the outline must be based on the subject information of the target domain; content that does not conform to the target domain theme is not included in the outline. The large language model encounters a significant discrepancy between the target text input and the overall topics covered in the descriptive information for the target domain and the extended domain. This is equivalent to the user inputting irrelevant document content, and the model directly outputs information that cannot be extracted from the target domain, thus terminating the human-computer interaction task.

[0104] It can be seen that by combining the target domain and the extended domain obtained by its supplementary extension, and constraining the description information of the domain in the general prompt information, the relevance of the output target domain outline content to the target domain can be strengthened, redundant content irrelevant to the target domain can be reduced, and the large language model can be inspired to generate a content outline that conforms to the thematic context of the target domain and suppress content in non-target domains.

[0105] In an optional embodiment, obtaining the enhanced prompt information corresponding to the target field includes: obtaining outline structure information, topic key information and output text length of the target field, and generating enhanced prompt information corresponding to the target field; wherein the outline structure information includes hierarchical identification, hierarchical depth and hierarchical structure.

[0106] The outline structure information is used to constrain the output format. The level identifier is used to identify the level of the directory. The level depth can be used to determine the number of levels in the directory. The level structure is used to determine the relationship between directories, such as the existence of a subordination relationship between directories.

[0107] In some embodiments, the directory is identified by #, and the number of hierarchical identifiers of the high-level directory is less than the number of hierarchical identifiers of the low-level directory. The number of hierarchical identifiers of adjacent high-level directories is less than the number of hierarchical identifiers of the low-level directories included in the high-level directory. For example, the first-level directory is identified by one #, the second-level directory is identified by two #s, and the third-level directory is identified by three #s. The hierarchical structure can be that the first-level directory includes the second-level directory, the second-level directory includes the third-level directory, and so on; and, the next adjacent line of the high-level directory is the low-level directory included in the high-level directory. When the low-level directory no longer includes the low-level directory, the next adjacent line of the low-level directory is the directory that is adjacent in directory order. The beginnings of sentences at the same level are aligned, and the indentation length of the beginnings of sentences of directories at different levels is adjusted according to the hierarchical depth. Usually, the indentation of the high-level directory is smaller than that of the low-level directory.

[0108] In an example, the extracted multi-level directory is as follows:

[0109] #Company Financial Performance

[0110] Q1 Financial Trends

[0111] ###First Quarter Highlights

[0112] #Business Update and Outlook

[0113] ##Content and Community Development

[0114] ##Commercialization progress and marginal expansion

[0115] ##Advertising Business Growth and Strategy

[0116] ##Gaming Business Outlook and Innovation

[0117] In the above example, the depth is 3, meaning the multi-level directory includes a first-level directory, a second-level directory, and a third-level directory. The indentation length of the first-level directory is smaller than that of the second-level directory, and the indentation length of the second-level directory is smaller than that of the third-level directory.

[0118] The key topic information is used to constrain the outline directory to be relevant to the target domain and target text, reducing the associations of the large language model. The output text length is used to constrain the length of the output text. In some embodiments, the key topic information can be used to directly output the content and prohibit the output of directory prompts. This can reduce the associations of the large language model and thus reduce the output of content irrelevant to the target text or target domain.

[0119] In one example, the enhanced tooltip looks like this:

[0120] #Level 1 directory (directly output content, prohibit output of level 1 directory prompt)

[0121] ##Level 2 subdirectory (directly output content, prohibit output of level 2 subdirectory prompt)

[0122] ##Level 2 subdirectory

[0123] #1 Level Directory

[0124] ##Level 2 subdirectory

[0125] ##Level 2 subdirectory

[0126] ##Level 2 subdirectory

[0127] Through experiments, the use of enhanced prompt information for constraints has achieved a 100% compliance rate in the test data with output instructions (mainly referring to the absence of redundant prompts and explanation information).

[0128] In addition, the enhanced prompt information can also include data enhancement information. For example, the outline directory must be fully discussed in the original text to increase the attention of the large language model to the target text and reduce hallucinations and associations.

[0129] It can be seen that by adding outline structure information, key topic information and output text length restrictions in the enhanced prompt information, it is possible to impose strong constraints on the outline content of the target field and improve the accuracy of the outline.

[0130] In an optional embodiment, after generating the target domain outline content based on the associated content corresponding to the directory of each directory information extraction task, it also includes: obtaining the positioning time corresponding to each sentence in the target text; and determining the positioning time corresponding to each associated content based on the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory.

[0131] Among them, the target text is the text obtained by speech recognition. The positioning time can refer to the starting time of the sentence in the speech. Specifically, the positioning time can be the time when the first word in the sentence is uttered in the speech. The sentences corresponding to the associated content in the target text can be associated. According to the positioning time corresponding to each sentence, the positioning time corresponding to the associated content is integrated. The positioning time of each sentence in the target text can be obtained by time positioning the sentence in the speech when the text of each sentence is obtained by speech recognition. In addition, the sentences corresponding to all the associated contents corresponding to the directory are the sentences corresponding to the directory. The sentences corresponding to all the directories included in the upper-level directory are the sentences corresponding to the upper-level directory. Therefore, the sentence range of each directory and the positioning time of each directory can be determined based on the order of the directories from the bottom to the top of the associated content.

[0132] In some embodiments, the speech-recognized text needs to be located based on the generated target domain outline content, to the content location corresponding to each outline directory, and / or to the time location in the speech.

[0133] It should be noted that the step of obtaining the positioning time corresponding to the associated content may not be implemented through a large language model.

[0134] In some embodiments, when the associated content information is extracted and the outline content of the target field in markdown format is generated, it is necessary to locate the original time of each specific content output of markdown (not the title) and return it to the server. This allows the product to click on each sentence of markdown output to display the corresponding original text and find the original time of the audio. Users can instantly browse to the location of each directory in the audio.

[0135] It can be seen that by obtaining the positioning time corresponding to each sentence in the target text and the associated content corresponding to the directory, matching the associated content with each sentence, and obtaining the positioning time corresponding to the directory, the demand for time positioning of the outline directory in the speech can be realized in the scenario of speech summary content generation, and time positioning based on precise associated content can improve the accuracy of positioning time.

[0136] In an optional embodiment, the method of determining the positioning time corresponding to each directory based on the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory includes: obtaining a sentence range corresponding to at least one associated content based on the associated content corresponding to each directory; for each associated content, calculating the hit probability of the associated content for each alternative sentence within the sentence range of the associated content; screening out the target sentence based on the hit probability of each alternative sentence; querying the positioning time corresponding to the target sentence based on the positioning time corresponding to each sentence in the target text; and determining the positioning time of the associated content based on the positioning time of the target sentence.

[0137] Compared to related content and directories, directories cover a wider range of content, making it difficult for users to quickly understand the original text of the directory from a larger range of content. Related content covers a smaller range of content, allowing users to quickly understand the original text of the related content from a smaller range of content. This makes it more efficient to locate related content.

[0138] Obtain the sentence range of the associated content in the target text. The sentence range includes at least one alternative sentence. The hit probability may refer to the probability of correlation between the associated content and the alternative sentence. The hit probability may be expressed as similarity. The target sentence may refer to the sentence that best represents the associated content among the alternative sentences. For each associated content, a target sentence is selected. The purpose of selecting a target sentence from the alternative sentences is to retain only one sentence to represent the associated content, which can achieve streamlined sentence positioning and accurate extraction of outline content.

[0139] It can be seen that by time locating the associated content and reducing the time positioning of the advanced directory, the effective time positioning data can be streamlined, and the positioning time of the most representative target sentence within the sentence range corresponding to the associated content can be screened out as the positioning time of the associated content. This can achieve accurate extraction of the positioning time and reduce the positioning time of redundant content, thereby achieving the refinement of the outline content.

[0140] In an optional embodiment, within the sentence scope of the associated content, the hit probability of the associated content for each alternative sentence within the sentence scope of the associated content is calculated, including: within the sentence scope of the associated content, calculating the hit probability of the associated content and the higher-level directory to which the associated content belongs for each alternative sentence within the sentence scope of the associated content.

[0141] The associated content is a directory obtained by subdividing the content of the higher-level directory to which the associated content belongs. Accordingly, the content of the higher-level directory is related to the associated content. The hit probability can be calculated by taking both the associated content and the higher-level directory as hit objects.

[0142] In some embodiments, the hit probability may refer to respectively calculating the text lengths of the alternative statement and two consecutive directories that hit. Optionally, the total hit length of the higher-level directory in the alternative statement and the total hit length of the associated content in the alternative statement are calculated. The weighted sum of the two hit total lengths is calculated as the hit probability. In some embodiments, the associated content is segmented to obtain at least one segmented text, each segmented text is matched with the alternative statement, and it is detected whether there are words that are hit by each segmented text in the alternative statement, and the lengths of the hit words are accumulated to obtain the total length of the segmented text of the associated content in the alternative statement. Among them, the hit words can be words that are the same as or have the same semantics as the segmented text. Similarly, by replacing the associated content with a higher-level directory, the total length of the segmented text of the higher-level directory in the alternative statement can be calculated.

[0143] In one example, the hit probability score is calculated using the following formula (1):

[0144] score=len_term_last+weight*len_term_target(1)

[0145] Where len_term_last is the total length of the segmented text of the higher-level directory in the candidate sentence hits, and len_term_target is the total length of the segmented text of the associated content in the candidate sentence hits. weight is the weight, and the experimental result is weight = 2.

[0146] It can be seen that by calculating the hit probability between the associated content and its higher-level directory and the same alternative statement as the hit probability between the associated content and the alternative statement, the relevant semantics of the associated content can be expanded. By calculating the hit probability based on the expanded semantics, the calculation accuracy of the hit probability of the alternative statement can be improved, thereby improving the accuracy of the alternative statement screening.

[0147] In an optional embodiment, the candidate sentences include: a single sentence and a continuous double sentence.

[0148] In fact, valid information related to the associated content may be simultaneously located in two adjacent sentences, so the two adjacent sentences can be used as an alternative sentence.

[0149] In an example, the statement range includes 10 statements. Each single sentence can be used as an alternative statement to obtain 10 alternative statements, and every two adjacent single sentences can be used as alternative statements to obtain 9 alternative statements. Accordingly, the statement range can determine 19 alternative statements.

[0150] Accordingly, the hit probability score of each single sentence is calculated using formula (1), and is used as the hit probability of the associated content for the single sentence. The penalty probability penalty of consecutive double sentences is calculated using the following formula (2):

[0151] penalty=(total_len_last+weight*total_len_target) / 6(2)

[0152] The hit probability score of each sentence in the continuous double sentence is calculated based on formula (1), minus the penalty probability calculated based on formula (2), to obtain the hit probability of the continuous double sentence, which is used as the hit probability of the associated content for the continuous double sentence.

[0153] It can be seen that by configuring the alternative sentences as single sentences or consecutive double sentences, different situations of the sentences where the valid information is located can be covered, so that one alternative sentence can completely cover the valid information, thereby screening the alternative sentences to obtain the target sentence, increasing the representativeness of the target sentence, and accurately locating the time of the related content.

[0154] In one example, the target text contains time information. The content of the target text is as follows:

[0155] 00:03:33A

[0156] 00:03:45B

[0157] 00:03:50C

[0158] 00:04:07D

[0159] 00:04:13E

[0160] 00:04:27F

[0161] The low-level directories for which positioning time needs to be obtained include the following:

[0162] ##Level 2 subdirectory

[0163] ###3 level subdirectory

[0164] -**Time**:XX

[0165] -**Reaction**: ZZ

[0166] -**Location**: YY

[0167] ##Level 2 subdirectory

[0168] In fact, -**time**:XX is the associated content. Among them, time and XX contain less information. You can first find the higher-level directory of this content, for example, the third-level subdirectory. Record XX as sent_target and the third-level subdirectory as sent_last.

[0169] Segment sent_target and sent_last, and remove information that does not contribute significantly to similarity calculations. For example, information that does not contribute significantly to similarity calculations includes punctuation and markdown symbols. The text set after segmentation is recorded as term_target, term_last.

[0170] The hit probability of each sentence A, sentence B, sentence C, sentence D, sentence E and sentence F.

[0171] For example, for a single sentence A, if "XX" appears in the original sentence A, the hit length is 2. If the third-level subdirectory appears in the original sentence A, the hit length is 3. Calculate the hit probability of term_target for single sentence A and the hit probability of term_last for single sentence A respectively, and get the hit probabilities of term_target and term_last for single sentence A.

[0172] For consecutive double sentences AB, BC, CD, DE and EF, calculate the hit probability of the double sentences.

[0173] For example, the penalty probability of consecutive double sentences AB is calculated. The hit probability of double sentences AB is the sum of the hit probability of single sentence A and the hit probability of single sentence B, minus the penalty probability of AB, to obtain the hit probability for consecutive double sentences AB.

[0174] Find the alternative sentence with the highest hit probability among single sentences A, single sentence B, single sentence C, single sentence D, single sentence E, single sentence F, AB, BC, CD, DE and EF, and use it as the target sentence. The positioning time of the target sentence is determined as the positioning time of the associated content.

[0175] In one example,

[0176] The low-level directories for which positioning time needs to be obtained include the following:

[0177] ##Level 2 subdirectory

[0178] ###3 level subdirectory

[0179] -**Time**:XX(00:03:50)

[0180] - **Reaction**: ZZ (00:04:07)

[0181] -**Location**: YY(00:04:27)

[0182] ##Level 2 subdirectory

[0183] In fact, during the calculation of positioning time, because the content to be retrieved has already clearly defined the retrieval time range (generated in the LLM process), the server's retrieval time can be controlled within 1ms, including the time for word segmentation and retrieval, which fully meets the requirements of real-time retrieval.

[0184] Figure 3 The present invention relates to a flowchart of another target text processing method disclosed in an embodiment of the present invention, which is further optimized and expanded based on the above technical solution and can be combined with the above optional implementation methods. The target prompt information corresponding to the target field and the target text are input into the large language model for processing to output text summary content, which is specifically as follows: the outline prompt information corresponding to the target field and the target text are input into the large language model for processing to generate the target field outline content; and the digital prompt information corresponding to the target field and the target text are input into the large language model for processing to generate the digital cluster summary content; and the digital prompt information corresponding to the target field is generated in a manner as follows: when the enhancement type includes contrast enhancement, the contrast subject enhancement extraction information of the target field is generated and added to the digital prompt information corresponding to the target field; when the enhancement type includes output description enhancement, the output description information of the target field is generated and added to the digital prompt information corresponding to the target field, wherein the output description information includes a description field and an attribute value field; when the enhancement type includes translation enhancement, the translation enhancement information of the target field is generated and added to the digital prompt information corresponding to the target field.

[0185] S301: Obtain target text.

[0186] S302: Acquire target prompt information corresponding to the target domain; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target domain.

[0187] S303: Input the outline prompt information corresponding to the target domain and the target text into a large language model for processing to generate the target domain outline content; and / or

[0188] S304: Input the digital prompt information corresponding to the target field and the target text into a large language model for processing to generate digital cluster summary content.

[0189] S305 , the digital prompt information corresponding to the target area is generated in the following manner: when the enhancement type includes contrast enhancement, contrast subject enhancement extraction information of the target area is generated, and added to the digital prompt information corresponding to the target area.

[0190] The contrast enhancement type can refer to the type of contrast data that needs to be enhanced in data comparison and analysis scenarios. Contrast subject enhancement extraction information can refer to information that extracts the main content of the target text from the previous and next sections. Numerical prompt information can refer to prompting the large language model to extract numerical summary content from the target text.

[0191] In one example, the target domain is finance. When year-on-year and / or quarter-on-quarter data descriptions appear in the target text, the main comparative information must be extracted from the context of the target text. For example, a 15% year-on-year increase must include the time frame for the year-on-year increase, such as year-on-year growth in 2024 or year-on-year growth in the third quarter.

[0192] S306: When the enhancement type includes output description enhancement, generate output description information of the target field and add it to the digital prompt information corresponding to the target field, wherein the output description information includes a description field and an attribute value field.

[0193] The enhanced output description type can refer to the need to limit (or constrain) the data output format and content in data analysis scenarios. Output description information can refer to extracting descriptive information from the target text. The description field is used to describe the data content, and the attribute value field is used to describe the data value.

[0194] In some embodiments, the name and value fields are used to express specific data content. The output description information requires that the data description be placed in the name field, and the specific numerical information be placed in the value field. During implementation, different few shots can be added sequentially to enhance the expressiveness of the numbers.

[0195] In an optional embodiment, the target field is the financial field, and the attribute value field includes: numbers and statistical units.

[0196] Numbers indicate numerical values. In one example, a few shots were used to fully extract the data's descriptive information and place it in the name field. The numerical value of the data was placed in the value field, which contained only the number and statistical unit, such as 15% or $1.3 million.

[0197] It can be seen that by configuring the attribute value field as numbers and statistical units, data information can be extracted reasonably and appropriately, and key information can be extracted while reducing redundant information.

[0198] S307: When the enhancement type includes translation enhancement, generate translation enhancement information of the target domain and add the translation enhancement information to the digital prompt information corresponding to the target domain.

[0199] The enhanced type of translation enhancement may refer to limiting the output format and output content of the language to be translated in the data analysis scenario. The translation enhancement information may refer to information about the target language of the translation.

[0200] In an optional embodiment, the target field is the financial field, and the translation enhancement information includes: content that the format of numbers is a common format of the language and / or correct enhancement prompt content for ambiguous statistical units.

[0201] The language-universal format may refer to a format that is common in multiple languages. For example, the language-universal format for numbers may refer to Arabic numerals.

[0202] In one example, the target text is in English, and the generated summary is in Chinese. Numbers in the English text need to be converted to their corresponding Arabic numerals, rather than their Chinese equivalents.

[0203] The correct reinforcement prompt content of the ambiguous statistical unit may refer to the content of the correct translation of the statistical unit that strengthens the translation ambiguity.

[0204] In one example, easily confused English units, such as billion, require a large language model to focus on whether the translation is correct. This is done by adding the correct reinforcement prompt content of the ambiguous statistical unit to the digital prompt: "**Billion is the unit of one billion and needs to be converted correctly**".

[0205] It can be seen that by limiting the direction of digital conversion in different languages, the accuracy of English parsing and the ability of expression can be improved, and by limiting the translation of ambiguous statistical units to strengthen sentences, the translation accuracy can be improved.

[0206] Existing large models with small number of parameters cannot reasonably express data description information by adjusting prompt, such as:

[0207] "name":"Year-on-year user growth rate",

[0208] "value":"130%"

[0209] This related content fails to effectively extract the data representation of the year-on-year growth rate. For example:

[0210] "name":"Number of users",

[0211] "value":"Year-on-year growth rate of 130% in 2023 compared to 2022"

[0212] This method is inappropriate in terms of data description and data value extraction. The reasonable extraction method should be:

[0213] "name":"Year-on-year growth rate of user numbers in 2023",

[0214] "value":"130%"

[0215] Through comparison and strengthening, the data representation of year-on-year growth can be effectively extracted.

[0216] Another issue with extracting data from large models with a small number of parameters is the poor parsing accuracy and expressiveness of English descriptions. For example, the original text reads: "seventy six point five billion dollars." The associated content for the large model is: "Seventy six and a half billion U.S. dollars." First, the data is incorrectly parsed; "billion" should be in the tens of billions, not hundreds of millions. Second, a better expression would be Arabic numerals, so the appropriate generation model would be: "67.5 billion U.S. dollars." By limiting the direction of English numerical conversion, the parsing accuracy and expressiveness of English descriptions can be improved. Furthermore, by limiting the translation of ambiguous English units, sentences can be strengthened and translation accuracy can be improved.

[0217] In some embodiments, the large language model outputs the extracted digital cluster summary content in JSON format as follows:

[0218] {"class":"UserData",

[0219] "data": [

[0220] {"name":"Year-on-year growth rate of identity authentication practitioners",

[0221] "value": "130%"}

[0222] {"name":"Average monthly subscription numbers of paid members in the first quarter",

[0223] "value":"1480"}

[0224] {"name":"Year-on-year growth rate of paid members in the first quarter",

[0225] "value": "0.8%"}

[0226] {"name":"Q1 paid membership growth rate",

[0227] "value":"4.1%"}]}

[0228] {"class":"Financial Data",

[0229] "data": [

[0230] {"name":"Marketing service revenue in the first quarter",

[0231] "value":"100 coins"}

[0232] {"name":"Percentage decrease in marketing service revenue year-on-year in the first quarter",

[0233] "value": "15.78%"}

[0234] {"name":"Year-on-year growth rate of advertising revenue in the first quarter",

[0235] "value":"4%"}

[0236] Q1 performance-based advertising business saw sequential growth

[0237] "value": "Unknown"}

[0238] {"name":"First quarter paid membership revenue",

[0239] "value": "4.45 coins"}]}

[0240] According to the technical solution disclosed in the present invention, by configuring the enhancement type as comparison enhancement, output description enhancement and translation enhancement, the digital prompt information is enhanced in different enhancement scenarios respectively, so that the digital prompt information can be enhanced from multiple dimensions, and the digital prompt information can be adaptively enhanced for data comparison scenarios, output scenarios and translation scenarios in different languages. The data description and numerical value can be effectively extracted, and the representation, language analysis, language expression and translation accuracy of the comparison data can be improved.

[0241] In one example, the generated text summary content is as follows Figure 4As shown in the figure, the left frame is the target text, which includes the role, speech time, and content. The right frame includes the target domain outline content and the digital cluster summary content. The target domain outline content includes multiple levels of directories, and the lowest level directory includes related content. The digital cluster summary content includes n clusters obtained by aggregation, and each class includes a description field and corresponding attribute value fields.

[0242] According to an embodiment of the present disclosure, Figure 5 This is a block diagram of a target text processing device in an embodiment of the present disclosure, which is suitable for extracting key content from a target text. The device is implemented using software and / or hardware and is specifically configured in an electronic device with certain data processing capabilities.

[0243] like Figure 5 The target text processing device 500 shown in the figure includes: a target text acquisition module 501, a prompt information determination module 502 and a text summary extraction module 503.

[0244] A target text acquisition module 501 is used to acquire the target text;

[0245] The prompt information determination module 502 is used to obtain target prompt information corresponding to the target field; the target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target field;

[0246] The text summary extraction module 503 is used to input the target prompt information corresponding to the target field and the target text into the large language model for processing and output text summary content. The target prompt information is used by the large language model to increase the attention of the strengthened constraint information to enhance the corresponding content in the text summary content.

[0247] According to the technical solution disclosed herein, target prompt information is generated based on the reinforcement constraint information corresponding to the reinforcement type of the target domain, so that the large language model pays more attention to the reinforcement constraint information, so that the large language model can perform data enhancement on the text summary content, reduce auditory hallucinations and comprehension difficulty in the text summary content, reduce content irrelevant to the target domain in the text summary content, increase content related to the target domain in the text summary content, and improve the accuracy of the text summary content in the target domain.

[0248] Optionally, the text summary extraction module includes at least one of the following:

[0249] an outline content generating unit, configured to input the outline prompt information corresponding to the target domain and the target text into a large language model for processing to generate the outline content of the target domain; and

[0250] The data content generating unit is used to input the digital prompt information corresponding to the target field and the target text into the large language model for processing to generate digital cluster summary content.

[0251] Optionally, the enhancement type includes: segmented processing;

[0252] The outline content generating unit includes:

[0253] A multi-level directory generating subunit, configured to generate a multi-level directory according to the first prompt information and the target text;

[0254] A directory task division subunit, configured to divide the multi-level directory into tasks and generate at least one directory information extraction task;

[0255] a directory content extraction subunit, configured to generate, for each directory information extraction task, associated content corresponding to the directory of the directory information extraction task according to the directory information extraction task, the second prompt information, and the target text;

[0256] The directory outline generation subunit is used to generate target domain outline content according to the associated content corresponding to the directory of each directory information extraction task.

[0257] Optionally, the directory task division subunit is specifically used to:

[0258] In the multi-level directory, obtain multiple high-level directories;

[0259] Divide each of the high-level directories into at least one task, where the upper limit of the number of divided tasks is a specified number;

[0260] For each of the high-level directories, the low-level directories associated with the high-level directory are divided into tasks divided by the high-level directory, and at least one directory information extraction task is generated.

[0261] Optionally, the directory task division subunit is specifically used to:

[0262] Obtaining the directory order between the high-level directories;

[0263] According to the order of each directory, each high-level directory is divided into at least one task, and the upper limit of the number of divided tasks is a specified number, wherein the number of directories included in the task of the high-level directory with an earlier directory order is less than the number of directories included in the task of the high-level directory with a later directory order.

[0264] Optionally, the multi-level directory includes: a first-level directory and a second-level directory.

[0265] Optionally, the enhancement type includes: character enhancement processing;

[0266] The device further includes: an outline prompt information generating module, wherein the outline prompt information generating module is configured to:

[0267] Obtaining general prompt information corresponding to the target field;

[0268] Obtaining enhanced prompt information corresponding to the target area;

[0269] Generate outline prompt information corresponding to the target field based on the general prompt information and the enhanced prompt information.

[0270] Optionally, the outline prompt information generating module is used to:

[0271] Obtaining the extended domain corresponding to the target domain;

[0272] Generate general prompt information corresponding to the target domain according to the description information of the target domain and the description information of the extended domain.

[0273] Optionally, the outline prompt information generating module is used to:

[0274] The outline structure information, key subject information and output text length of the target domain are obtained, and enhanced prompt information corresponding to the target domain is generated; wherein the outline structure information includes a level identifier, a level depth and a level structure.

[0275] Optionally, the target text processing device further includes: a digital prompt information generation module, wherein the digital prompt information generation module is configured to:

[0276] When the enhancement type includes contrast enhancement, generating contrast subject enhancement extraction information of the target area and adding the information to digital prompt information corresponding to the target area;

[0277] When the enhancement type includes output description enhancement, generating output description information of the target field and adding it to the digital prompt information corresponding to the target field, wherein the output description information includes a description field and an attribute value field;

[0278] When the enhancement type includes translation enhancement, translation enhancement information of the target domain is generated and added to the digital prompt information corresponding to the target domain.

[0279] Optionally, the target field is the financial field, and the attribute value field includes: numbers and statistical units.

[0280] Optionally, the target field is the financial field, and the translation enhancement information includes: content that the format of numbers is a common format of the language and / or correct enhancement prompt content for ambiguous statistical units.

[0281] Optionally, the target text processing device further includes:

[0282] A positioning time acquisition module is used to obtain the positioning time corresponding to each sentence in the target text after generating the target domain outline content based on the associated content corresponding to the directory of each directory information extraction task;

[0283] The directory time positioning module is used to determine the positioning time corresponding to each directory according to the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory.

[0284] Optionally, the directory time positioning module includes:

[0285] a statement range determining unit, configured to obtain, based on the associated contents corresponding to each of the directories, a statement range corresponding to at least one associated content;

[0286] a hit probability determination unit, configured to calculate, for each of the associated contents, within a sentence range of the associated content, a hit probability of the associated content for each candidate sentence within the sentence range of the associated content;

[0287] A hit statement screening unit, configured to screen out a target statement based on the hit probability of each candidate statement;

[0288] A positioning time query unit, configured to query the positioning time corresponding to the target sentence based on the positioning time corresponding to each sentence in the target text;

[0289] The positioning time determining unit is configured to determine the positioning time of the associated content according to the positioning time of the target sentence.

[0290] Optionally, the hit probability determination unit includes:

[0291] The association probability calculation subunit is used to calculate the hit probability of each candidate sentence in the sentence range of the associated content and the higher-level directory to which the associated content belongs within the sentence range of the associated content.

[0292] Optionally, the alternative sentences include: single sentences and consecutive double sentences.

[0293] The above-mentioned target text processing device can execute the target text processing method provided by any embodiment of the present disclosure, and has corresponding functional modules and beneficial effects for executing the target text processing method.

[0294] In the technical solutions disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0295] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0296] Figure 6 A schematic area diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0297] like Figure 6 As shown, the device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the instructions of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0298] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0299] The computing unit 601 can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the target text processing method. For example, in some embodiments, the target text processing method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the target text processing method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the target text processing method by any other suitable means (e.g., by means of firmware).

[0300] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard objects (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0301] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / instructions specified in the flow chart and / or area diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0302] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0303] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0304] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0305] A computer system may include a client and a server. The client and server are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0306] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0307] Cloud computing refers to a technology system that provides network access to elastically scalable shared pools of physical or virtual resources. These resources can include servers, instruction sets, networks, software, applications, and storage devices, and can be deployed and managed on-demand in a self-service manner. Cloud computing technology provides efficient and powerful data processing capabilities for the application of technologies such as artificial intelligence and blockchain, as well as for model training.

[0308] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided by this disclosure can be achieved. This is not a limitation herein.

[0309] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A target text processing method, comprising: Get the target text; Get target prompt information corresponding to the target area; The target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target field; Inputting the target prompt information corresponding to the target field and the target text into a large language model for processing, and outputting text summary content, wherein the target prompt information is used by the large language model to increase attention to the strengthened constraint information to enhance the corresponding content in the text summary content; The step of inputting the target prompt information corresponding to the target domain and the target text into a large language model for processing and outputting text summary content includes at least one of the following: Inputting the outline prompt information corresponding to the target domain and the target text into the large language model for processing to generate the target domain outline content; as well as Inputting the digital prompt information corresponding to the target field and the target text into a large language model for processing to generate digital cluster summary content; The enhancement type includes: segmentation processing; the enhancement type refers to the dimension of data enhancement, and the enhancement constraint information is the descriptive information corresponding to the enhancement type, which is used to enhance the text comprehension ability and the extraction ability of the target domain content; The step of inputting the outline prompt information corresponding to the target domain and the target text into a large language model for processing to generate the target domain outline content includes: Generate a multi-level directory based on the first prompt information and the target text; Dividing the multi-level directory into tasks to generate at least one directory information extraction task; For each of the directory information extraction tasks, generating associated content corresponding to the directory of the directory information extraction task according to the directory information extraction task, the second prompt information and the target text; Generate target domain outline content based on the associated content corresponding to the directory of each directory information extraction task.

2. The method according to claim 1, wherein The step of dividing the multi-level directory into tasks to generate at least one directory information extraction task includes: In the multi-level directory, obtain multiple high-level directories; Divide each of the high-level directories into at least one task, where the upper limit of the number of divided tasks is a specified number; For each of the high-level directories, the low-level directories associated with the high-level directory are divided into tasks divided by the high-level directory, and at least one directory information extraction task is generated.

3. The method according to claim 2, wherein: Dividing each of the high-level directories into at least one task includes: Obtaining the directory order between the high-level directories; According to the order of each directory, each high-level directory is divided into at least one task, wherein the number of directories included in the task of the high-level directory with an earlier directory order is smaller than the number of directories included in the task of the high-level directory with a later directory order.

4. The method according to claim 1 or 2, wherein: The multi-level directory includes: a first-level directory and a second-level directory.

5. The method according to claim 1, wherein The enhancement types include: character enhancement processing; the outline prompt information corresponding to the target area is generated in the following way: Obtaining general prompt information corresponding to the target field; Obtaining enhanced prompt information corresponding to the target area; Generate outline prompt information corresponding to the target field based on the general prompt information and the enhanced prompt information.

6. The method according to claim 5, wherein: The obtaining of general prompt information corresponding to the target field includes: Obtaining the extended domain corresponding to the target domain; Generate general prompt information corresponding to the target domain according to the description information of the target domain and the description information of the extended domain.

7. The method according to claim 5, wherein: The obtaining of enhanced prompt information corresponding to the target field includes: The outline structure information, key subject information and output text length of the target domain are obtained, and enhanced prompt information corresponding to the target domain is generated; wherein the outline structure information includes a level identifier, a level depth and a level structure.

8. The method according to claim 1, wherein The digital prompt information corresponding to the target field is generated in the following way: When the enhancement type includes contrast enhancement, generating contrast subject enhancement extraction information of the target area and adding the information to digital prompt information corresponding to the target area; When the enhancement type includes output description enhancement, generating output description information of the target field and adding it to the digital prompt information corresponding to the target field, wherein the output description information includes a description field and an attribute value field; When the enhancement type includes translation enhancement, translation enhancement information of the target domain is generated and added to the digital prompt information corresponding to the target domain.

9. The method according to claim 8, wherein The target field is the financial field, and the attribute value field includes: numbers and statistical units.

10. The method according to claim 8, wherein The target field is the financial field, and the translation enhancement information includes: content that the format of numbers is a common format of the language and / or correct enhancement prompt content for ambiguous statistical units.

11. The method according to claim 1, after generating target domain outline content based on the associated content corresponding to the directory of each directory information extraction task, further comprising: Obtaining the location time corresponding to each sentence in the target text; The positioning time corresponding to each associated content is determined according to the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory.

12. The method according to claim 11, wherein The determining, based on the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory, the positioning time corresponding to each associated content includes: According to the associated content corresponding to each of the directories, obtaining a statement range corresponding to at least one associated content; For each of the related contents, within the sentence range of the related content, calculating the hit probability of the related content for each candidate sentence within the sentence range of the related content; Filtering out the target sentence according to the hit probability of each candidate sentence; According to the positioning time corresponding to each sentence in the target text, query the positioning time corresponding to the target sentence; The positioning time of the associated content is determined according to the positioning time of the target sentence.

13. The method according to claim 12, wherein: The calculating, within the sentence range of the associated content, the hit probability of the associated content for each candidate sentence within the sentence range of the associated content includes: Within the sentence range of the associated content, the hit probability of each candidate sentence within the sentence range of the associated content and the higher-level directory to which the associated content belongs is calculated.

14. The method according to claim 12 or 13, wherein: The alternative sentences include: single sentences and consecutive double sentences.

15. A target text processing device, comprising: A target text acquisition module is used to acquire the target text; A prompt information determination module is used to obtain target prompt information corresponding to the target field; The target prompt information is generated according to the reinforcement constraint information corresponding to the reinforcement type in the target field; a text summary extraction module, configured to input the target prompt information corresponding to the target domain and the target text into a large language model for processing, and output text summary content, wherein the target prompt information is used by the large language model to increase attention to the strengthened constraint information to enhance the corresponding content in the text summary content; The text summary extraction module includes at least one of the following: An outline content generating unit, configured to input the outline prompt information corresponding to the target domain and the target text into a large language model for processing to generate the outline content of the target domain; as well as A data content generating unit, configured to input the digital prompt information corresponding to the target field and the target text into a large language model for processing to generate digital cluster summary content; The enhancement type includes: segmentation processing; the enhancement type refers to the dimension of data enhancement, and the enhancement constraint information is the descriptive information corresponding to the enhancement type, which is used to enhance the text comprehension ability and the extraction ability of the target domain content; The outline content generating unit includes: A multi-level directory generating subunit, configured to generate a multi-level directory according to the first prompt information and the target text; A directory task division subunit, configured to divide the multi-level directory into tasks and generate at least one directory information extraction task; a directory content extraction subunit, configured to generate, for each directory information extraction task, associated content corresponding to the directory of the directory information extraction task according to the directory information extraction task, the second prompt information, and the target text; The directory outline generation subunit is used to generate target domain outline content according to the associated content corresponding to the directory of each directory information extraction task.

16. The device according to claim 15, wherein The directory task division subunit is specifically used to: In the multi-level directory, obtain multiple high-level directories; Divide each of the high-level directories into at least one task, where the upper limit of the number of divided tasks is a specified number; For each of the high-level directories, the low-level directories associated with the high-level directory are divided into tasks divided by the high-level directory, and at least one directory information extraction task is generated.

17. The device according to claim 16, wherein The directory task division subunit is specifically used to: Obtaining the directory order between the high-level directories; According to the order of each directory, each high-level directory is divided into at least one task, wherein the number of directories included in the task of the high-level directory with an earlier directory order is smaller than the number of directories included in the task of the high-level directory with a later directory order.

18. The device according to claim 15 or 16, wherein The multi-level directory includes: a first-level directory and a second-level directory.

19. The device according to claim 15, wherein The enhancement types include: role strengthening processing; The device further includes: an outline prompt information generating module, wherein the outline prompt information generating module is configured to: Obtaining general prompt information corresponding to the target field; Obtaining enhanced prompt information corresponding to the target area; Generate outline prompt information corresponding to the target field based on the general prompt information and the enhanced prompt information.

20. The device according to claim 19, wherein The outline prompt information generating module is used to: Obtaining the extended domain corresponding to the target domain; Generate general prompt information corresponding to the target domain according to the description information of the target domain and the description information of the extended domain.

21. The apparatus according to claim 19, wherein The outline prompt information generating module is used to: The outline structure information, key subject information and output text length of the target domain are obtained, and enhanced prompt information corresponding to the target domain is generated; wherein the outline structure information includes a level identifier, a level depth and a level structure.

22. The apparatus according to claim 15, further comprising: A digital prompt information generation module, wherein the digital prompt information generation module is used to: When the enhancement type includes contrast enhancement, generating contrast subject enhancement extraction information of the target area and adding the information to digital prompt information corresponding to the target area; When the enhancement type includes output description enhancement, generating output description information of the target field and adding it to the digital prompt information corresponding to the target field, wherein the output description information includes a description field and an attribute value field; When the enhancement type includes translation enhancement, translation enhancement information of the target domain is generated and added to the digital prompt information corresponding to the target domain.

23. The device according to claim 22, wherein The target field is the financial field, and the attribute value field includes: numbers and statistical units.

24. The apparatus according to claim 22, wherein The target field is the financial field, and the translation enhancement information includes: content that the format of numbers is a common format of the language and / or correct enhancement prompt content for ambiguous statistical units.

25. The apparatus of claim 15, further comprising: A positioning time acquisition module is used to obtain the positioning time corresponding to each sentence in the target text after generating the target domain outline content based on the associated content corresponding to the directory of each directory information extraction task; The directory time positioning module is used to determine the positioning time corresponding to each associated content according to the positioning time corresponding to each sentence in the target text and the associated content corresponding to each directory.

26. The device according to claim 25, wherein The directory time positioning module includes: a statement range determining unit, configured to obtain, based on the associated contents corresponding to each of the directories, a statement range corresponding to at least one associated content; a hit probability determination unit, configured to calculate, for each of the associated contents, within a sentence range of the associated content, a hit probability of the associated content for each candidate sentence within the sentence range of the associated content; A hit statement screening unit, configured to screen out a target statement based on the hit probability of each candidate statement; A positioning time query unit, configured to query the positioning time corresponding to the target sentence based on the positioning time corresponding to each sentence in the target text; The positioning time determining unit is configured to determine the positioning time of the associated content according to the positioning time of the target sentence.

27. The device according to claim 26, wherein The hit probability determination unit includes: The association probability calculation subunit is used to calculate the hit probability of each candidate sentence in the sentence range of the associated content and the higher-level directory to which the associated content belongs within the sentence range of the associated content.

28. The device according to claim 26 or 27, wherein The alternative sentences include: single sentences and consecutive double sentences.

29. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any target text processing method according to claims 1-14.

30. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the target text processing method according to any one of claims 1-14.

31. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the target text processing method according to any one of claims 1 to 14.

Citation Information

Patent Citations

  • Audio and video data processing method and device, equipment and storage medium

    CN112860939A

  • Voice control method and device, electronic equipment and readable storage medium

    CN116705018A