Problem-oriented Intelligent Summarization Method, Device, Medium and Product for Enterprise Services

By identifying and expanding the subject words in business problems, filtering and clustering document paragraphs, and using large language models to generate problem-oriented intelligent summary of enterprise services, the problems of inflexible expression and insufficient content integrity in the existing technology are solved, and efficient professional text data processing is achieved.

CN119938907BActive Publication Date: 2025-07-04SUYIDA (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510041374.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2025-07-04
Estimated Expiration
2045-01-10

AI Technical Summary

Technical Problem

When processing highly professional text data, the existing intelligent summary technology is not flexible enough, the content is complete and credible, and lacks business background information, which affects reading and comprehension efficiency.

Method used

Natural language processing technology is used to identify the subject words in business problems, and expand them based on business graphs, main data information or business question responses, filter relevant document paragraphs, generate abstracts through aggregation hierarchical clustering algorithm and large language model, and perform semantic vector representation and visual display.

Benefits of technology

The generated abstract results enrich the business knowledge background information, retain the completeness of the text content, improve reading and comprehension efficiency, and flexible expression forms, which are suitable for the reading needs of corporate users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119938907B_ABST
    Figure CN119938907B_ABST
Patent Text Reader

Abstract

The present invention discloses a problem-oriented intelligent summarization method, device, medium and product for enterprise services, belonging to the technical field of data processing, including: obtaining business problems, using natural language processing technology to identify the subject words in the business problems, expanding the subject words based on the business graph or master data information to obtain subject word expansion terms; screening the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list; screening the documents in the first document list to obtain a first paragraph; clustering the first paragraph based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, and generating a summary result and a total summary, performing semantic vector representation on the summary result and the total summary to generate a first semantic vector, sorting the semantic vectors, generating a visual interface and displaying it, which can process texts with strong professionalism, has a flexible expression form, and ensures the integrity of the content and the accuracy of the summary, facilitating user reading.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a problem-oriented intelligent summarization method, device, medium and product for enterprise services. Background Art

[0002] At present, the Internet has entered the big data era, and text data has become one of the most important resources in enterprises and organizations. Existing intelligent summarization technologies can refine and compress text data, improving the efficiency of readers' understanding of text data. Intelligent summarization technologies are mainly divided into two categories: extractive summarization technology and generative summarization technology. Among them, extractive summarization technology splits text data into multiple components according to specific rules, scores the importance of each part or uses a binary classification method to judge whether it is important, and retains the parts with higher importance scores or matching important tags in the text data as the summary result; early generative summarization technology adopted the method of predefined templates, extracted key information elements from text data, and filled the key information elements into the predefined templates. However, with the development and application of generative artificial intelligence technology, summary results can be automatically generated through the Transformer architecture.

[0003] However, the foregoing technologies have the following deficiencies: the expression forms of extractive summarization technology and early generative summarization technology are not flexible enough and the readability is poor; for the generative summarization technology based on the Transformer architecture, it is easy to lose text data, and the integrity and credibility of the content are insufficient; all text data is analyzed and processed without discrimination, and the relevance between the generated summary results and the actual information needs is poor; the summary results lack a content theme structure, affecting the reading and understanding efficiency; for highly professional text data, the lack of understanding of business background information and knowledge affects the importance scoring or the judgment of important information. Summary of the Invention

[0004] The purpose of the present invention is to provide a problem-oriented intelligent summarization method, device, medium and product for enterprise services to solve the above problems existing in the prior art.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] In a first aspect, the present invention provides a problem-oriented intelligent summarization method for enterprise services, including:

[0007] Obtain a business problem, use natural language processing technology to identify the subject words in the business problem, and expand the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information or business problem answers;

[0008] Screen the titles and bodies of the documents in the database based on the subject terms and the extended terms of the subject terms to obtain a first document list, where the first document list contains at least one document whose title or body contains the subject terms and / or the extended terms of the subject terms;

[0009] Screen the documents in the first document list according to the relevance of business problems to obtain at least one first paragraph;

[0010] Cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and use the large language model again to extract the summary result to generate a total summary corresponding to the summary result. Perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector;

[0011] Sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visualization interface, and display the visualization interface.

[0012] In a possible design, the subject terms include nouns or noun phrases; correspondingly, when using natural language processing technology to identify the subject terms in the business problem and expand the subject terms based on the business graph to obtain the extended terms of the subject terms, the method includes:

[0013] Use word segmentation technology and syntactic analysis technology to identify the business problem to obtain the subject terms;

[0014] Match the business graph with the subject terms to obtain the entity nodes corresponding to the subject terms;

[0015] Obtain the first-order adjacent nodes in the business graph according to the subject terms and the entity nodes, where the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes;

[0016] According to the semantic relevance, perform semantic similarity sorting on the first-order adjacent nodes to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarities in the semantic similarity sequence as the extended terms of the subject terms.

[0017] In a possible design, according to the semantic relevance, perform semantic similarity sorting on the first-order adjacent nodes to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarities in the semantic similarity sequence as the extended terms of the subject terms, including:

[0018] Extract the corresponding triples in the business graph based on the subject terms, where the triples are (X, ri, Yi), where X is the subject term, ri is the relationship name of the triple, and Yi is the first-order adjacent node;

[0019] Concatenate the triples into a string, calculate the semantic similarity between the business problem and the string, and sort them in descending order of semantic similarity to obtain a semantic similarity sequence. Among them, concatenating the triples into a string is Ti = concate(X, ri, Yi), where Ti is the string and concate() is the concatenation function;

[0020] Select the first-order adjacent nodes corresponding to the top three semantic similarities in the semantic similarity sequence as the topic word expansion terms.

[0021] In a possible design, when using natural language processing technology to identify the topic words in the business problem and expanding the topic words based on the business problem reply to obtain the topic word expansion terms, the method includes:

[0022] Input the business problem into a large language model to obtain a business problem reply, and expand the business problem based on the business problem reply to obtain an expanded business problem;

[0023] Identify the topic word expansion terms based on the word segmentation dictionary and natural language processing technology for the expanded business problem.

[0024] In a possible design, when using natural language processing technology to identify the topic words in the business problem and expanding the topic words based on the master data information to obtain the topic word expansion terms, the method includes:

[0025] Obtain the master data information and input the master data information into the word segmentation dictionary;

[0026] Identify the master data information of the business problem based on the word segmentation dictionary and natural language processing technology to obtain the topic words;

[0027] Expand the topic words with other master data information corresponding to the master data in the business graph to obtain the topic word expansion terms.

[0028] In a possible design, screening the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph, including:

[0029] Extract the natural paragraphs containing the topic words or topic word expansion terms as the first paragraphs;

[0030] Select the natural paragraph with the most occurrences of the topic words or topic word expansion terms in the document as the positive feature typical paragraph;

[0031] Convert the natural paragraphs in the document based on the Embedding function to obtain at least one semantic vector, and calculate the semantic vector similarity between at least one semantic vector and the semantic vector of the positive feature typical paragraph;

[0032] Select the natural paragraph with the lowest semantic vector similarity in the document, which does not contain the topic word and does not contain the extended terms of the topic word as the negative feature typical paragraph;

[0033] Based on the positive feature typical paragraph and the negative feature typical paragraph, compare the semantic distances of the natural paragraphs that do not contain the topic word or the extended terms of the topic word. If the semantic distance between the natural paragraph and the positive feature typical paragraph is greater than the semantic distance between the natural paragraph and the negative feature typical paragraph, extract it as the first paragraph.

[0034] In a possible design, cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, including:

[0035] Extract all the first paragraphs of at least one document as subclasses and select the cluster center;

[0036] If the number of subclasses is greater than the preset number threshold, cluster the first paragraphs based on the agglomerative hierarchical clustering algorithm until the number of clustering subclasses reaches the preset number threshold, and output the clustering subclasses as the clustering result; if the number of subclasses is less than or equal to the preset number threshold, output the subclasses as the clustering result.

[0037] In a second aspect, the present invention provides a problem-oriented intelligent summarization device for enterprise services, including:

[0038] An identification and expansion unit, configured to obtain a business problem, identify the topic word in the business problem by using natural language processing technology, and expand the topic word based on business information to obtain extended terms of the topic word, where the business information includes a business graph, master data information, or a business problem reply;

[0039] A first screening unit, configured to screen the titles and texts of the document list in the database based on the topic word and the extended terms of the topic word to obtain a first document list, where the first document list contains at least one document whose title or text contains the topic word and / or the extended terms of the topic word;

[0040] A second screening unit, configured to screen the documents in the first document list according to the business problem relevance to obtain at least one first paragraph;

[0041] A clustering and extraction unit, configured to cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and then use the large language model to extract the summary result to generate a total summary corresponding to the summary result, and perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector;

[0042] A sorting and display unit is used to sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visual interface, and display the visual interface.

[0043] In a third aspect, the present invention provides a computer-readable storage medium, on which instructions are stored. When the instructions run on a computer, they execute the problem-oriented intelligent summarization method for enterprise services as described in any one of the above.

[0044] In a fourth aspect, the present invention provides a computer program product containing instructions. When the instructions run on a computer, the computer is made to execute the problem-oriented intelligent summarization method for enterprise services as described in any one of the above.

[0045] Beneficial effects: The present invention discloses a problem-oriented intelligent summarization method, device, medium, and product for enterprise services. It uses natural language processing technology to perform word segmentation on business problems, identify the topic words in the business problems, and expand the topic words based on business graphs, master data information, or business problem answers. It is targeted at enterprise users, has rich business knowledge background information, and can process text data with strong professionalism. It analyzes and screens all documents to ensure the integrity of the content, and only retains the paragraphs containing the topic words or the expanded terms of the topic words. It extracts the generated summary results again to generate a total summary, and performs visual processing on the total summary and the summary results, and displays the visual interface. The expression form is flexible, which is convenient for users to read and understand. Description of the Drawings

[0046] Figure 1 It is a flowchart of the problem-oriented intelligent summarization method for enterprise services provided by an embodiment of the present invention;

[0047] Figure 2 It is a block diagram schematic of the problem-oriented intelligent summarization device for enterprise services provided by an embodiment of the present invention. Detailed Embodiments

[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0049] It should be understood that although terms such as first and second may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, the first unit may be referred to as the second unit, and similarly, the second unit may be referred to as the first unit, without departing from the scope of the exemplary embodiments of the present invention.

[0050] It should be understood that for the term "and / or" that may appear herein, it is merely an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, B exists alone, and A and B exist simultaneously; for the term " / and" that may appear herein, it is a description of another association object relationship, indicating that two relationships may exist. For example, A / and B may represent: A exists alone, and A and B exist alone; in addition, for the character " / " that may appear herein, generally it represents that the associated objects before and after are in an "or" relationship.

[0051] Embodiment:

[0052] As Figure 1 shown, this embodiment provides a problem-oriented intelligent summarization method for enterprise services, which may but is not limited to including the following steps.

[0053] S1. Obtain business problems, use natural language processing technology to identify the subject words in the business problems, and expand the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information, or business problem answers;

[0054] Among them, the subject words include nouns or noun phrases in the business problems; the business graph is an enterprise-level knowledge graph that records data such as business terms, work processes, technical standards, and rules and regulations of relevant enterprises; the master data information is the name, code, abbreviation, alias, and English name of business objects.

[0055] Specifically, in step S1, when using natural language processing technology to identify the subject words in the business problems and expanding the subject words based on the business graph to obtain subject word expansion terms, the method includes:

[0056] S101. Use word segmentation technology and syntactic analysis technology to identify the business problems to obtain subject words;

[0057] Preferably, the word segmentation technology includes but is not limited to regular expressions, Hidden Markov Model (HMM), Word2Vec model, and GloVe model (Global Vectors for Word Representation); the syntactic analysis technology includes but is not limited to semantic association analysis, semantic field analysis, and semantic template matching.

[0058] S102. Match the business graph with the subject word to obtain the entity nodes corresponding to the subject word;

[0059] S103. Obtain the first-order adjacent nodes in the business graph according to the subject word and the entity nodes, where the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes;

[0060] In the embodiment of the present application, for example, if the entity node is an energy monitoring system, there are positive and negative relationships of the entity node in the business graph. For example, the positive relationship is "energy monitoring system" -> "function" -> "monitor abnormal power consumption"; the negative relationship is "charging pile" -> "access" -> "energy monitoring system", then the first-order forward adjacent node is to monitor abnormal power consumption, and the first-order reverse adjacent node is the charging pile.

[0061] S104. According to the semantic relevance, perform semantic similarity sorting on the first-order adjacent nodes to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarity from the semantic similarity sequence as the subject word expansion terms.

[0062] Among them, in step S104, according to the semantic relevance, perform semantic similarity sorting on the first-order adjacent nodes to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarity from the semantic similarity sequence as the subject word expansion terms, including:

[0063] S1041. Extract the corresponding triples in the business graph based on the subject word, where the triple is (X, ri, Yi), where X is the subject word, ri is the relationship name of the triple, and Yi is the first-order adjacent node;

[0064] S1042. Concatenate the triples into a string, calculate the semantic similarity between the business problem and the string, and sort them from high to low according to the semantic similarity to obtain a semantic similarity sequence, where concatenating the triples into a string is Ti = concate(X, ri, Yi), where Ti is the string and concate() is the concatenation function;

[0065] S1043. Select the first three first-order adjacent nodes corresponding to the semantic similarity in the semantic similarity sequence as the subject word expansion terms.

[0066] In a possible design, in step S1, when using natural language processing technology to identify the subject words in the business problem and expanding the subject words based on the business problem response to obtain subject word expansion terms, the method includes:

[0067] S101. Input the business problem into a large language model to obtain a business problem response, and expand the business problem based on the business problem response to obtain an expanded business problem;

[0068] S102. Identify the subject word expansion terms based on the word segmentation dictionary and natural language processing technology for the expanded business problem.

[0069] In a possible design, in step S1, when using natural language processing technology to identify the subject words in the business problem and expanding the subject words based on the master data information to obtain subject word expansion terms, the method includes:

[0070] S101. Obtain the master data information and input the master data information into the word segmentation dictionary;

[0071] S102. Identify the master data information of the business problem based on the word segmentation dictionary and natural language processing technology to obtain the subject words;

[0072] S103. Expand the subject words using other master data information corresponding to the master data in the business graph to obtain subject word expansion terms.

[0073] In the embodiments of the present application, the master data includes, but is not limited to, supplier data, product data, equipment data, and material data, and the master data information includes the name, code, abbreviation, alias, and English name, etc. in the master data.

[0074] S2. Screen the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list, where the first document list contains at least one document whose title or text contains the subject words and / or subject word expansion terms;

[0075] S3. Screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph;

[0076] Specifically, in step S3, screening the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph includes:

[0077] S301. Extract the natural paragraphs containing the subject words or subject word expansion terms as the first paragraphs;

[0078] S302. Select the natural paragraph with the most occurrences of the subject words or the extended terms of the subject words in the document as the positive characteristic typical paragraph;

[0079] S303. Based on the Embedding function, transform the natural paragraphs in the document to obtain at least one semantic vector, and calculate the semantic vector similarity between at least one semantic vector and the semantic vector of the positive characteristic typical paragraph;

[0080] The Embedding function, that is, the embedding function, maps other types of entities, such as sentences, documents, items, or people, etc., to another numerical vector space through a specific method.

[0081] S304. Select the natural paragraph with the lowest semantic vector similarity in the document and that does not contain the subject words and does not contain the extended terms of the subject words as the negative characteristic typical paragraph;

[0082] S305. Based on the positive characteristic typical paragraph and the negative characteristic typical paragraph, compare the semantic distances of the natural paragraphs that do not contain the subject words or the extended terms of the subject words. If the semantic distance between the natural paragraph and the positive characteristic typical paragraph is greater than the semantic distance between the natural paragraph and the negative characteristic typical paragraph, then extract it as the first paragraph.

[0083] S4. Based on the agglomerative hierarchical clustering algorithm, cluster the first paragraphs in at least one document in the first document list to obtain a clustering result. Use a large language model to extract from the clustering result to generate at least one summary result, and then use the large language model again to extract from the summary result to generate the total summary corresponding to the summary result. Represent the summary result and the total summary corresponding to the summary result in semantic vectors to generate the first semantic vector;

[0084] Specifically, in step S4, clustering the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result includes:

[0085] S401. Extract all the first paragraphs of at least one document as subclasses and select the clustering center;

[0086] Among them, the selection method of the clustering center is the prior art and can be selected according to the usage needs, which will not be elaborated here.

[0087] S402. If the number of subclasses is greater than the preset number threshold, then perform clustering processing on the first paragraphs based on the agglomerative hierarchical clustering algorithm until the number of clustering subclasses reaches the preset number threshold, and output the clustering subclasses as the clustering result; if the number of subclasses is less than or equal to the preset number threshold, then output the subclasses as the clustering result.

[0088] Preferably, the preset number threshold can be 5.

[0089] S5. Sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visualization interface, and display the visualization interface.

[0090] In the implementation process of this embodiment, first, obtain a business problem, use natural language processing technology to identify the subject words in the business problem, and expand the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information, or business problem answers; screen the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list, where the first document list contains at least one document whose title or text contains the subject words and / or subject word expansion terms; screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and use the large language model again to extract the summary result to generate a total summary corresponding to the summary result, perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector; sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visualization interface, and display the visualization interface. It has rich business knowledge background information, can process text data with strong professionalism, is targeted at enterprise users, and has a flexible expression form, which is convenient for users to read and understand.

[0091] As Figure 2 shown, the second aspect of this embodiment provides a problem-oriented intelligent summary device for enterprise services, including:

[0092] An identification and expansion unit, configured to obtain a business problem, use natural language processing technology to identify the subject words in the business problem, and expand the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information, or business problem answers;

[0093] A first screening unit, configured to screen the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list, where the first document list contains at least one document whose title or text contains the subject words and / or subject word expansion terms;

[0094] A second screening unit, configured to screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph;

[0095] A clustering extraction unit is used to cluster the first paragraphs in at least one document based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and use the large language model again to extract the summary result to generate a total summary corresponding to the summary result, and perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector;

[0096] A sorting and display unit sorts according to the semantic vector similarity between the business problem and the first semantic vector, generates a visual interface, and displays the visual interface.

[0097] In summary, the problem-oriented intelligent summary device for enterprise services provided in this embodiment obtains a business problem through an identification and expansion unit, uses natural language processing technology to identify the subject words in the business problem, and expands the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information, or business problem answers; the first screening unit screens the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list; the second screening unit screens the documents in the first document list according to the business problem relevance to obtain at least one first paragraph; the clustering extraction unit clusters the first paragraphs in at least one document in the first document list using the agglomerative hierarchical clustering algorithm to obtain a clustering result, uses a large language model to extract the clustering result to generate at least one summary result, and uses the large language model again to extract the summary result to generate a total summary corresponding to the summary result, and performs semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector; the sorting and display unit sorts according to the semantic vector similarity between the business problem and the first semantic vector, generates a visual interface, and displays the visual interface, analyzes and screens all documents, ensures the integrity of the content, and only retains the paragraphs with high similarity to the subject words or subject word expansion terms, improving the accuracy of summary generation.

[0098] In the third aspect of this embodiment, a computer-readable storage medium is provided. Instructions are stored on the computer-readable storage medium, and when the instructions run on a computer, they are used to execute the problem-oriented intelligent summary method for enterprise services as described in the first aspect of the embodiment. Among them, the computer-readable storage medium refers to a carrier for storing data, and may include, but is not limited to, computer-readable storage media such as floppy disks, optical discs, hard disks, flash memories, USB flash drives, and / or memory sticks. The computer may be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices.

[0099] For the working process, working details and technical effects of the aforementioned computer-readable storage medium provided in the third aspect of this embodiment, reference may be made to the intelligent abstract generation method described in the first aspect, which will not be elaborated herein.

[0100] In the fourth aspect of this embodiment, a computer program product is provided, including a computer program or instruction, which, when executed by a computer, is used to implement the problem-oriented intelligent abstract method for enterprise services described in the first aspect of the embodiment.

[0101] For the working process, working details and technical effects of the aforementioned computer program product provided in the fourth aspect of this embodiment, reference may be made to the intelligent distributed data scraping method described in the first aspect, which will not be elaborated herein.

[0102] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A problem-oriented intelligent summarization method for enterprise services, characterized in that, Including: Obtain a business problem, use natural language processing technology to identify the subject words in the business problem, and expand the subject words based on business information to obtain subject word expansion terms, where the business information includes a business graph, master data information, or business problem answers; Screen the titles and texts of the document list in the database based on the subject words and subject word expansion terms to obtain a first document list, where the first document list contains at least one document whose title or text contains the subject words and / or subject word expansion terms; Screen the documents in the first document list according to the relevance of the business problem to obtain at least one first paragraph; Cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and then use the large language model to extract the summary result to generate a total summary corresponding to the summary result, and perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector; Sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visualization interface, and display the visualization interface.

2. The problem-oriented intelligent abstract method for enterprise services according to claim 1, characterized in that The subject words include nouns or noun phrases; correspondingly, when using natural language processing technology to identify the subject words in the business problem and expand the subject words based on the business graph to obtain subject word expansion terms, the method includes: Use word segmentation technology and syntactic analysis technology to identify the business problem to obtain subject words; Match the business graph with the subject words to obtain entity nodes corresponding to the subject words; Obtain the first-order adjacent nodes in the business graph according to the subject words and entity nodes, where the first-order adjacent nodes include first-order forward adjacent nodes and first-order reverse adjacent nodes; Sort the first-order adjacent nodes according to semantic relevance to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarity in the semantic similarity sequence as the subject word expansion terms.

3. The problem-oriented intelligent abstract method for enterprise services according to claim 2, wherein Sort the first-order adjacent nodes according to semantic relevance to obtain a semantic similarity sequence, and select the first three first-order adjacent nodes corresponding to the semantic similarity in the semantic similarity sequence as the subject word expansion terms, including: Extract the corresponding triples in the business graph based on the subject words, where the triples are (X, ri, Yi), where X is the subject word, ri is the relationship name of the triple, and Yi is the first-order adjacent node; Concatenate the triples into a string, calculate the semantic similarity between the business problem and the string, and sort them from high to low according to the semantic similarity to obtain a semantic similarity sequence, where concatenating the triples into a string is Ti = concate(X, ri, Yi), where Ti is the string and concate() is the concatenation function; Select the first three first-order adjacent nodes corresponding to the semantic similarity in the semantic similarity sequence as the subject word expansion terms.

4. The problem-oriented intelligent abstract method for enterprise services according to claim 1, characterized in that When using natural language processing technology to identify the subject words in the business problem and expanding the subject words based on the business problem response to obtain the subject word expansion terms, the method includes: Input the business problem into a large language model to obtain a business problem response, and expand the business problem based on the business problem response to obtain an expanded business problem; Based on the word segmentation dictionary and natural language processing technology, identify the expanded business problem to obtain the subject word expansion terms.

5. The problem-oriented intelligent abstract method for enterprise services according to claim 1, characterized in that When using natural language processing technology to identify the subject words in the business problem and expanding the subject words based on the master data information to obtain the subject word expansion terms, the method includes: Obtain the master data information and input the master data information into the word segmentation dictionary; Based on the word segmentation dictionary and natural language processing technology, identify the master data information in the business problem to obtain the subject words; Use other master data information corresponding to the master data in the business graph to expand the subject words to obtain the subject word expansion terms.

6. The problem-oriented intelligent abstract method for enterprise services according to claim 1, characterized in that, Screen the documents in the first document list according to the business problem relevance to obtain at least one first paragraph, including: Extract the natural paragraphs containing the subject words or the subject word expansion terms as the first paragraphs; Select the natural paragraph with the most occurrences of the subject words or the subject word expansion terms in the document as the positive feature typical paragraph; Based on the Embedding function, transform the natural paragraphs in the document to obtain at least one semantic vector, and calculate the semantic vector similarity between at least one semantic vector and the semantic vector of the positive feature typical paragraph; Select the natural paragraph with the lowest semantic vector similarity in the document and that does not contain the subject words and does not contain the subject word expansion terms as the negative feature typical paragraph; Based on the positive feature typical paragraph and the negative feature typical paragraph, compare the semantic distances of the natural paragraphs that do not contain the subject words or the subject word expansion terms. If the semantic distance between the natural paragraph and the positive feature typical paragraph is greater than the semantic distance between the natural paragraph and the negative feature typical paragraph, then extract it as the first paragraph.

7. The problem-oriented intelligent abstract method for enterprise services according to claim 1, characterized in that Cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, including: Extract all the first paragraphs of at least one document in the first document list as subclasses and select the clustering center; If the number of subclasses is greater than the preset number threshold, then perform clustering processing on the first paragraphs based on the agglomerative hierarchical clustering algorithm until the number of clustering subclasses reaches the preset number threshold, and output the clustering subclasses as the clustering result; if the number of subclasses is less than or equal to the preset number threshold, then output the subclasses as the clustering result.

8. A problem-oriented intelligent abstract device for enterprise services, characterized in that, Include: An identification and expansion unit, configured to obtain a business problem, use natural language processing technology to identify the subject words in the business problem, and expand the subject words based on the business information to obtain the subject word expansion terms, where the business information includes a business graph, master data information, or a business problem response; A first screening unit, configured to screen the titles and texts of the document list in the database based on the subject words and the subject word expansion terms to obtain a first document list, where the first document list contains at least one document whose title or text contains the subject words and / or the subject word expansion terms; A second screening unit, configured to screen documents in the first document list according to business problem relevance to obtain at least one first paragraph; A clustering extraction unit, configured to cluster the first paragraphs in at least one document in the first document list based on the agglomerative hierarchical clustering algorithm to obtain a clustering result, use a large language model to extract the clustering result to generate at least one summary result, and again use the large language model to extract the summary result to generate a total summary corresponding to the summary result, perform semantic vector representation on the summary result and the total summary corresponding to the summary result to generate a first semantic vector; A sorting and display unit, configured to sort according to the semantic vector similarity between the business problem and the first semantic vector, generate a visualization interface, and display the visualization interface.

9. A computer-readable storage medium, characterized in that, Instructions are stored on the computer-readable storage medium, and when the instructions are run on a computer, the problem-oriented intelligent summarization method for enterprise services described in any one of claims 1 to 7 is executed.

10. A computer program product, comprising a computer program or instructions, characterized in that, The computer program or the instructions, when executed by a computer, implement the problem-oriented intelligent summarization method for enterprise services described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Document processing method and system based on natural language and knowledge graph

    CN116501875A

  • Abstract generation method, abstract generation device, electronic equipment and readable storage medium

    CN118093859A